Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Paras Dahal, Anton Bakhtin, Taco Cohen, Rob Fergus, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal and colleagues at Meta Superintelligence Labs introduce agentic meta-reasoning, an inference-time harness in which a controller decides what work to assign, so that choosing the next step becomes a structured reasoning process separate from the task work.
Ask this paper
Controller and workers. Workers do task-level computation; the controller runs a four-stage cycle (update state, propose next computations, value them under the remaining budget, dispatch with selected earlier outputs). Controller calls count against the same budget as worker calls.
Compact state, persistent memory. Between decisions the controller keeps only a short account of the run; full worker outputs stay in memory and are retrieved when needed, so the controller's context does not grow with run length.
ProgramBench. With GPT-5.5 the harness reaches 71.5% mean test-pass rate against 58.0% for Codex and 63.7% for a Direct Control Agent using the same workers; with Opus 4.8 it reaches 67.2% against 65.5% for Claude Code.
Other benchmarks. On IMO ProofBench-Advanced, ARC-AGI-2 and LongCoT-mini it gains 3.6 to 4.2 points over direct control averaged over Gemini 3.1 Pro, GPT-5.5 and Opus 4.8, and it is ahead in all 12 matched comparisons at the main budgets.
Scaling and limits. Gains keep growing over the tested budget range where direct control plateaus, but controller overhead lowers scores at small budgets. The authors note the comparison tests the whole design, not each stage separately.
Abstract
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.