AutoLab
Free while signed in. Answers cite the passages they came from.

Can frontier models actually grind on a hard engineering problem the way a good researcher does? AutoLab is a benchmark for ultra long-horizon, closed-loop optimization built to answer that. It contains 36 realistic, expert-curated tasks across four domains: system optimization, puzzle and challenge, model development, and CUDA kernel optimization. Each task hands the agent a correct but deliberately suboptimal baseline and asks it to improve within a strict wall-clock budget.
Persistence beats a strong start: The dominant predictor of final performance is not the quality of the initial solution but the agent's persistence in iterative refinement. Models that keep probing and improving win, regardless of where they began.
Most models quit early: While Claude Opus 4.6 shows strong long-horizon optimization, most frontier models, including several proprietary ones, either terminate prematurely or burn their budget with minimal progress.
Time awareness is the gap: The results point to time-awareness and sustained iteration, not raw single-shot capability, as the missing ingredient for truly capable long-horizon agents.
Why it matters: Day-one benchmarks reward clever first attempts, but real research and engineering reward stamina. AutoLab measures the thing that actually separates agents on multi-hour tasks, and the benchmark, harness, and task artifacts are open-sourced.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack