🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 10, 2026
Agents

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

First page
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
The curator’s take

Yuan Tian, Bing Hu and colleagues (independent researchers with UC Berkeley, Purdue and Stanford) cross seed and self-evolved harnesses with base and LoRA-trained weights to decide when an agent should be improved through its harness and when through its weights.

Ask this paper

Key points
01

Failure composition. Failed trajectories are labelled by the first signal that fires, separating process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor).

02

Harness evolution. On DeepPlanning, a self-evolving harness lifts held-out score from 0.16 to 0.30 for Qwen3.5-4B and from 0.32 to 0.44 for Qwen3.5-9B, and raises 4B held-out delivery from 55% to 90%, mostly by repairing process failures.

03

Weight training. LoRA adapters trained on evolved-harness trajectories add +0.13 on held-out tasks under the original harness for both sizes; on 9B the adapter alone matches the full evolution line and cuts content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model.

04

Transfer. The loop gives +0.09 on 117 unseen WebArena-Lite tasks, where adapters add nothing on top of the harness.

05

Rule. Read the failure composition to choose between harness and weights, then read what the accepted harness edits changed to decide which gains to train into the model.

Abstract

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.

Every Monday
Get next week’s papers.
Subscribe on Substack