🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Reinforcement Learning · Training

Harness-R1

First page
Harness-R1
Paper summary

Agents accumulate interaction trajectories during deployment and then leave them unused, because their behavior stays fixed. Those trajectories can improve the harness that constructs context, mediates tools, validates actions, and recovers execution, and this work makes that editing a learned capability.

Ask this paper

Key points
01

A dedicated harness engineer: A separate 9B model converts batches of target-agent failures into validated executable patches across the runtime lifecycle, initialized with cold-start supervised fine-tuning and then trained online with group-relative policy optimization.

02

The target stays frozen: Fresh same-batch reruns of the frozen target supply outcome rewards, so training updates only the engineer and the agent being repaired holds still under the reward signal.

03

It works before and after tuning the target: Across WebShop, ALFWorld, and DBBench, vanilla Qwen3.5-9B goes from 44.3% to 53.6%, and after the target itself is fine-tuned a target-specific engineer lifts the average further from 59.2% to 64.2%.

04

Why it matters: If you run agents in production you already have the training data, and because the gains hold on both sides of target fine-tuning, the paper points toward co-evolving the harness engineer and the agent it repairs.

Every Monday
Get next week’s papers.
Subscribe on Substack