Mixture of Self-Improving Branches For Agent Harness Optimization

Haoyu Dong, Zihao Lin, Lizhu Zhang, Zhuokai Zhao and colleagues at Meta (with Duke and UC Davis) extend Meta-Harness-style harness optimization by splitting the search into branches, each with its own evolving development subset and proposal policy, and then routing each new input to one branch's best harness.
Ask this paper
Problem with one trajectory. Meta-Harness keeps a fixed development set and proposal policy, so all edits follow one search path and can settle in a local optimum.
Branch objectives. Each branch keeps the development cases its leading harnesses solve more often than other branches do, drops cases that every branch already solves, and rewrites its proposal policy from its own search history.
Router. Before execution a router picks one development-selected branch head per input, combining complementary harnesses without looking at test outcomes.
Results. Relative gains over Meta-Harness of 34.8% on Olympiad-level math, 11.6% on Terminal-Bench 2.0 and 3.8% on SWE-bench Lite, with harness selection and router configuration using development data only.
Abstract
Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.