🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation

Long-form factuality in LLMs

Free while signed in. Answers cite the passages they came from.

First page
Long-form factuality in LLMs
The curator’s take

Google DeepMind introduces LongFact and SAFE, a prompt set and automated evaluator for judging whether the long-form answers of modern LLMs are actually factual.

Key points
01

LongFact prompts: LongFact contains thousands of questions across 38 topics designed to elicit multi-paragraph, fact-dense responses rather than short answers.

02

SAFE evaluator: SAFE decomposes each response into atomic claims, issues Google Search queries for each, and uses an LLM agent to judge whether each claim is supported.

03

Superhuman agreement: SAFE matches crowdsourced human annotators on 72% of claims and, on disagreements, wins 76% of the time - while being roughly 20x cheaper per annotation.

04

Bigger = more factual: Across 13 evaluated models, larger LLMs generally produce more factually accurate long-form answers, and the team reports F1@K (precision balanced against recall at K claims) as a new long-form factuality metric.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack