🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Evaluation · Training

Rehearse

First page
Rehearse
Paper summary

Autoresearch loops propose changes, run full training jobs, and keep whatever improves the metric. Their efficiency depends on judging, before spending a run, whether a proposed modification is likely to work, and this paper studies how that judgment holds up over a trajectory.

Ask this paper

Key points
01

The capability is real at first: On 296 same-baseline modification pairs from 39 paper-derived tasks with outcomes hidden, an LLM judge given rationales but no prior-attempt history reaches 79.5% accuracy where strict consensus returns a verdict.

02

The confidence cliff: Across the full 366-pair benchmark, selective accuracy falls from 82.8% to 56.9% as successful changes accumulate while the judge stays just as willing to decide, and in public AutoSOTA logs the fraction of helpful modifications drops from 70% in the first two iterations to 43% by iteration six.

03

Propose, predict, execute: Rehearse is a small loop change shipped as a lightweight skill. Propose several ideas, compare them before execution, run the most promising, and judge against a focused memory of similar past attempts and their outcomes.

04

Why it matters: Late selective accuracy recovers to 83.5%, and across 4,000 budgeted training runs on nanochat, image classification, and time-series forecasting the endpoint improves under the same budget, which makes this a cheap patch for anyone running a self-improving loop.

Every Monday
Get next week’s papers.
Subscribe on Substack