🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Long-form factuality in LLMs

First page
Long-form factuality in LLMs
Paper summary

Google DeepMind introduces LongFact and SAFE, a prompt set and automated evaluator for judging whether the long-form answers of modern LLMs are actually factual.

Ask this paper

Key points
01

LongFact prompts: LongFact contains thousands of questions across 38 topics designed to elicit multi-paragraph, fact-dense responses rather than short answers.

02

SAFE evaluator: SAFE decomposes each response into atomic claims, issues Google Search queries for each, and uses an LLM agent to judge whether each claim is supported.

03

Superhuman agreement: SAFE matches crowdsourced human annotators on 72% of claims and, on disagreements, wins 76% of the time - while being roughly 20x cheaper per annotation.

04

Bigger = more factual: Across 13 evaluated models, larger LLMs generally produce more factually accurate long-form answers, and the team reports F1@K (precision balanced against recall at K claims) as a new long-form factuality metric.

Every Monday
Get next week’s papers.
Subscribe on Substack