🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning · Evaluation

Sample More Reflect Less

Free while signed in. Answers cite the passages they came from.

First page
Sample More Reflect Less
The curator’s take

Methods that make a model criticize and rewrite its own answer nearly all generate far more text than a single chain of thought. Since generating more text raises accuracy on its own, a reported gain leaves open whether the method's idea is what helped. This paper reruns the comparison as a designed experiment.

Key points
01

Every token counted: Seven methods, open models at 1.5B, 3B, and 7B, two math benchmarks with 150 questions each, and every generated token counted including critiques, reflections, debate turns, and checking, with each method compared against repeated sampling at its own measured cost.

02

No reliable win anywhere: All 36 comparisons are paired by question with bootstrap intervals and multiplicity correction, and repeated sampling holds up against every method at equal cost in every setting.

03

Self-inspection is the failure mode: Ten comparisons come back reliably worse and every one of them is a method where the model inspects its own output, with all 18 self-inspection comparisons negative. Reflexion as published never triggered its own retry on the smallest model because it judged itself correct every time.

04

Why it matters: Adding a critique step is the default reflex when an agent loop underperforms, and this study runs the comparison with paired bootstrap intervals and multiplicity correction, which the earlier point-estimate comparison lacked.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack