RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Chen-Yu Lee, Tomas Pfister and colleagues at Google Cloud AI Research, with UNC Chapel Hill, Stanford and WashU, show that automated harness evolution overfits its training tasks and add regularization to both the edit proposer and the selector.
Ask this paper
The problem. Iteratively proposing and selecting harness edits produces large in-distribution gains that shrink or disappear on out-of-distribution benchmarks.
Regularized proposer. A temporally annealed budget limits how many edits one candidate can bundle, and evolution history is used to push toward unexplored trajectories.
Critic and pruner. The critic screens out benchmark-specific proposals, and the pruner removes changes that are too small, too expensive or no longer useful.
Results. Across eight benchmarks in coding, agentic workspace and engineering design, RRSI gains up to 14.1 points on the evolution split and up to 4.7 points on five out-of-distribution benchmarks.
Cheaper harness. The final harness uses 30% fewer policy tokens than unregularized evolution; code is at github.com/google-research/rrsi.
Abstract
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.