🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Reasoning · Agents

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

First page
Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
The curator’s take

Ankur Samanta, Kaveh Hassani and Anirudh Goyal at Meta AI, with Yonathan Efroni (Tel Aviv) and Paul Sajda (Columbia), introduce MIRA, a research-agent architecture in which an outer meta-reasoner decides what to investigate next and a fresh executor carries out each investigation, and train that outer policy with RL.

Ask this paper

Key points
01

Architecture. The meta-reasoner curates context from a persistent research record and writes a work order for the next investigation, or ends the episode. A new inner executor runs each work order, so execution becomes part of the transition between meta-level decisions.

02

Without training. MIRA improves long-horizon inference and uses extra compute better on IMOProofBench theorem proving and on open-ended architecture research with Residual Matrix Transformers and Loop Transformers.

03

Decision-level critic. A generative critic trained at work-order boundaries forecasts remaining return better than token-level alternatives, and cross-environment pretraining improves forecasting and adaptation.

04

MIRA-AC. A single generative actor-critic trained on the model's own proxy hill-climbing signals improves gold-evaluation performance in all four autoresearch environments (symbolic regression, physics discovery, Bayesian causal discovery, CPU architecture research), and the actor transfers to an environment it was not trained on.

Abstract

Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model's own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.

Every Monday
Get next week’s papers.
Subscribe on Substack