Harness-R1
Free while signed in. Answers cite the passages they came from.

Agents accumulate interaction trajectories during deployment and then leave them unused, because their behavior stays fixed. Those trajectories can improve the harness that constructs context, mediates tools, validates actions, and recovers execution, and this work makes that editing a learned capability.
A dedicated harness engineer: A separate 9B model converts batches of target-agent failures into validated executable patches across the runtime lifecycle, initialized with cold-start supervised fine-tuning and then trained online with group-relative policy optimization.
The target stays frozen: Fresh same-batch reruns of the frozen target supply outcome rewards, so training updates only the engineer and the agent being repaired holds still under the reward signal.
It works before and after tuning the target: Across WebShop, ALFWorld, and DBBench, vanilla Qwen3.5-9B goes from 44.3% to 53.6%, and after the target itself is fine-tuned a target-specific engineer lifts the average further from 59.2% to 64.2%.
Why it matters: If you run agents in production you already have the training data, and because the gains hold on both sides of target fine-tuning, the paper points toward co-evolving the harness engineer and the agent it repairs.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack