Long-form factuality in LLMs

Google DeepMind introduces LongFact and SAFE, a prompt set and automated evaluator for judging whether the long-form answers of modern LLMs are actually factual.
Ask this paper
LongFact prompts: LongFact contains thousands of questions across 38 topics designed to elicit multi-paragraph, fact-dense responses rather than short answers.
SAFE evaluator: SAFE decomposes each response into atomic claims, issues Google Search queries for each, and uses an LLM agent to judge whether each claim is supported.
Superhuman agreement: SAFE matches crowdsourced human annotators on 72% of claims and, on disagreements, wins 76% of the time - while being roughly 20x cheaper per annotation.
Bigger = more factual: Across 13 evaluated models, larger LLMs generally produce more factually accurate long-form answers, and the team reports F1@K (precision balanced against recall at K claims) as a new long-form factuality metric.