🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning · Evaluation

Robust Evaluation of Reasoning

Free while signed in. Answers cite the passages they came from.

First page
Robust Evaluation of Reasoning
The curator’s take

The paper introduces functional benchmarks that parameterize reasoning problems so the same structural question can be re-instantiated with fresh surface forms, then uses them to expose a large "reasoning gap" in frontier LLMs.

Key points
01

Functional benchmarks: MATH() is built as a functional variant of the MATH benchmark, turning each item into a template that can generate many semantically equivalent but textually different instances.

02

Large reasoning gap: On functional variants, SoTA models drop between 58.35% and 80.31% relative to the static benchmark, suggesting much of headline performance is memorization or surface pattern matching.

03

Prompting helps: More sophisticated prompting strategies shrink (but do not eliminate) the gap, implying that eliciting the right reasoning traces is partially possible without changing the model.

04

Evaluation lesson: Static benchmarks overstate reasoning ability; functionalization is proposed as a cheap, generalizable way to stress-test new models.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack