Rehearse

Autoresearch loops propose changes, run full training jobs, and keep whatever improves the metric. Their efficiency depends on judging, before spending a run, whether a proposed modification is likely to work, and this paper studies how that judgment holds up over a trajectory.
Ask this paper
The capability is real at first: On 296 same-baseline modification pairs from 39 paper-derived tasks with outcomes hidden, an LLM judge given rationales but no prior-attempt history reaches 79.5% accuracy where strict consensus returns a verdict.
The confidence cliff: Across the full 366-pair benchmark, selective accuracy falls from 82.8% to 56.9% as successful changes accumulate while the judge stays just as willing to decide, and in public AutoSOTA logs the fraction of helpful modifications drops from 70% in the first two iterations to 43% by iteration six.
Propose, predict, execute: Rehearse is a small loop change shipped as a lightweight skill. Propose several ideas, compare them before execution, run the most promising, and judge against a focused memory of similar past attempts and their outcomes.
Why it matters: Late selective accuracy recovers to 83.5%, and across 4,000 budgeted training runs on nanochat, image classification, and time-series forecasting the endpoint improves under the same budget, which makes this a cheap patch for anyone running a self-improving loop.