🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Evaluation · Training

Rehearse

Free while signed in. Answers cite the passages they came from.

First page
Rehearse
The curator’s take

Autoresearch loops propose changes, run full training jobs, and keep whatever improves the metric. Their efficiency depends on judging, before spending a run, whether a proposed modification is likely to work, and this paper studies how that judgment holds up over a trajectory.

Key points
01

The capability is real at first: On 296 same-baseline modification pairs from 39 paper-derived tasks with outcomes hidden, an LLM judge given rationales but no prior-attempt history reaches 79.5% accuracy where strict consensus returns a verdict.

02

The confidence cliff: Across the full 366-pair benchmark, selective accuracy falls from 82.8% to 56.9% as successful changes accumulate while the judge stays just as willing to decide, and in public AutoSOTA logs the fraction of helpful modifications drops from 70% in the first two iterations to 43% by iteration six.

03

Propose, predict, execute: Rehearse is a small loop change shipped as a lightweight skill. Propose several ideas, compare them before execution, run the most promising, and judge against a focused memory of similar past attempts and their outcomes.

04

Why it matters: Late selective accuracy recovers to 83.5%, and across 4,000 budgeted training runs on nanochat, image classification, and time-series forecasting the endpoint improves under the same budget, which makes this a cheap patch for anyone running a self-improving loop.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack