🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Evaluation

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

First page
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
The curator’s take

Zongxia Li, Yucheng Shi, Zhongzhi Li and colleagues at Tencent HY LLM Frontier, with the University of Maryland and others, propose Recursive Self-Rewrite (RSR), which collects successful solutions under several specialized harnesses and rewrites them into training trajectories for one general harness.

Ask this paper

Key points
01

Discovery. Three harnesses (Terminus 2, StateM and Recursive Self-Reflect Terminus) solve complementary tasks; their union solves 759 of 3K curated terminal tasks, 34.3% more than the best single harness.

02

Rewrite pipeline. The same Qwen-3.8-27B model acts as planner (turns a solution into a runbook), critic (removes verifier and solution leakage) and executor (re-solves the task under Terminus 2 in a fresh sandbox). Rejection sampling yields about 10K verified trajectories.

03

Results. Training on rewritten trajectories raises pass@3 from 57.0% to 74.2% on Terminal-Bench 2, from 39.0% to 63.0% on Terminal-Bench Hard and from 1.5% to 9.1% on Terminal-Bench 4, and beats direct SFT on all five benchmarks (+20.8 points on TB2).

04

Why rewrite. Raw trajectories from specialized harnesses contain controller interventions that are unavailable at deployment; rewriting keeps the planning, checking and recovery steps while removing harness-specific conventions.

Abstract

Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.

Every Monday
Get next week’s papers.
Subscribe on Substack