GAMUT
Free while signed in. Answers cite the passages they came from.

Most factuality evaluation measures precision, whether the claims in an answer are correct. This Meta AI work targets the harder and mostly ignored half, completeness, meaning whether an answer covers everything it should, and packages it as the GAMUT benchmark.
Completeness is structured: The facts a complete answer should contain rarely form a flat list, since they involve open-ended sets where coverage matters, ordered processes, and relationships among facts that independent boolean checks cannot capture.
Two-level meta-rubrics: A structured meta-rubric encodes the organization and importance of required content, then compiles mechanically into a flat checklist of binary, machine-gradable items that an LLM judge can score reliably, keeping rich structure while inheriting low-variance grading.
Grounded and verified: The benchmark holds 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert annotators, and a text-only variant is released for models without vision.
Why it matters: Across 14 frontier and open-weight models the benchmark stays genuinely hard, with a best score of 58.7% from Gemini 3.1 Pro, while remaining highly discriminative and robust to the choice of judge.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack