🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Evaluation

Robust Evaluation of Reasoning

First page
Robust Evaluation of Reasoning
Paper summary

The paper introduces functional benchmarks that parameterize reasoning problems so the same structural question can be re-instantiated with fresh surface forms, then uses them to expose a large "reasoning gap" in frontier LLMs.

Ask this paper

Key points
01

Functional benchmarks: MATH() is built as a functional variant of the MATH benchmark, turning each item into a template that can generate many semantically equivalent but textually different instances.

02

Large reasoning gap: On functional variants, SoTA models drop between 58.35% and 80.31% relative to the static benchmark, suggesting much of headline performance is memorization or surface pattern matching.

03

Prompting helps: More sophisticated prompting strategies shrink (but do not eliminate) the gap, implying that eliciting the right reasoning traces is partially possible without changing the model.

04

Evaluation lesson: Static benchmarks overstate reasoning ability; functionalization is proposed as a cheap, generalizable way to stress-test new models.

Every Monday
Get next week’s papers.
Subscribe on Substack