🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Retrieval

FreshLLMs (FreshQA)

First page
FreshLLMs (FreshQA)
Paper summary

Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

Ask this paper

Key points
01

Dynamic QA benchmark: Continuously refreshes questions so models can't memorize answers - a direct response to the contamination concerns plaguing static benchmarks.

02

Four question categories: Covers never-changing, slow-changing, fast-changing, and false-premise questions, stressing different aspects of freshness handling.

03

Reveals freshness gap: Shows that LLMs without search augmentation answer fast-changing questions poorly, while retrieval-augmented models close most of the gap.

04

FreshPrompt: Proposes FreshPrompt, a simple search-augmented prompting strategy that substantially boosts LLM performance on time-sensitive questions.

Every Monday
Get next week’s papers.
Subscribe on Substack