🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Evaluation · Reasoning

Harness Evolution, Rethought

Free while signed in. Answers cite the passages they came from.

First page
Harness Evolution, Rethought
The curator’s take

Automatic harness evolution is what many teams now use to squeeze more out of agents, but the reported gains might not be coming from the harness at all. This paper argues that harness evolution is itself a search procedure and must be compared against simple search baselines under matched budgets.

Key points
01

A fairer comparison: Because harness evolution repeatedly evaluates and revises candidates using task feedback, it should be benchmarked against task-level search under the same feedback and inference budgets, not against a single static harness.

02

The gains do not hold up: On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, evolved harnesses fall to 67.4, below the 68.2 baseline, while plain parallel sampling reaches 72.3 and harness scaling reaches 71.8.

03

Weak generalization: Beyond underperforming simple test-time scaling, the evolved harnesses transfer poorly, undercutting the assumption that a searched configuration captures something durable.

04

Why it matters: The result is a caution for anyone banking on self-evolving harnesses, and a call for evaluation protocols that separate genuine harness benefit from the effect of simply spending more compute on search.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack