Robust Evaluation of Reasoning

The paper introduces functional benchmarks that parameterize reasoning problems so the same structural question can be re-instantiated with fresh surface forms, then uses them to expose a large "reasoning gap" in frontier LLMs.
Ask this paper
Functional benchmarks: MATH() is built as a functional variant of the MATH benchmark, turning each item into a template that can generate many semantically equivalent but textually different instances.
Large reasoning gap: On functional variants, SoTA models drop between 58.35% and 80.31% relative to the static benchmark, suggesting much of headline performance is memorization or surface pattern matching.
Prompting helps: More sophisticated prompting strategies shrink (but do not eliminate) the gap, implying that eliciting the right reasoning traces is partially possible without changing the model.
Evaluation lesson: Static benchmarks overstate reasoning ability; functionalization is proposed as a cheap, generalizable way to stress-test new models.