SelfSearch: Reward-Free Search for Self-Improving Agents

Jungwoo Yang, Injin Kong and Yohan Jo at Seoul National University introduce SelfSearch, in which coding agents rewrite their own instructions, tools and procedures using records of earlier self-modification episodes, with no downstream task reward during the search.
Ask this paper
Method. Each episode record captures the reasoning, tool actions and outcomes of a previous modification attempt. The modified agent becomes the next improver, so both task skill and self-modification skill can improve.
Gains without reward. Population-mean success rises over the initial agent in all six model-benchmark settings, with individual agents gaining up to 11.2 points on Terminal-Bench 2.1.
Efficiency. On SWE-bench Multilingual, one evolved agent gains 5.0 points and cuts execution cost by 38.5% on tasks both versions solve.
Cost. With $4.03 of search cost, SelfSearch produces a harness that solves 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, matching Codex, the top harness in a public nine-harness comparison under the same settings.
What the agents build. Code diffs show agents writing reusable tools, including a trajectory reader that supports later self-improvement, and fixing failures in their own tool interactions.
Abstract
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.