LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

Euntae Choi, Sumin Song and Sungjoo Yoo at Seoul National University present LiteEvo, a harness-evolution loop that starts from a benchmark-neutral harness, tells its optimizer nothing about the benchmark, costs about $8 to $13 per run, and produces harnesses that also help on held-out tasks.
Ask this paper
Motivation. Prior harness optimizers such as HarnessX spend 100 to 175 million meta-agent tokens per benchmark, start from handcrafted harnesses, and report gains only on the tasks they evolved on.
Method. Each benchmark starts from the same minimal harness. Workers propose changes to prompts, action and observation processing and loop settings; a composer combines them, and a candidate is accepted only if it beats the current best by a noise-aware margin.
AppWorld result. A frozen Qwen3.5-9B solves 0% of AppWorld test tasks from the task text alone. With a LiteEvo harness its pass@2 rises to 66.1% when evolved on the test tasks and to 71.4% when evolved on separate training tasks.
Cost and coverage. LiteEvo improves pass@2 in all nine benchmark and regime cells at $8.21 to $12.92 per run, matching or beating a HarnessX reproduction at 13x lower mean API cost.
Held-out tasks. Harnesses evolved on training tasks stay within 2.2 points of transductive runs on ALFWorld and AppWorld, and score higher on WebShop (70.0 vs 62.7) and SWE-bench Verified (66.7 vs 63.2).
Abstract
An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark. We propose LiteEvo, a lightweight harness-evolution algorithm whose tool-free meta-agents mine agent trajectories for reusable components, curate them into a versioned library, and compose each round's harness from it, starting every benchmark from the same neutral harness and never naming the benchmark. Evolving on the graded tasks of five agentic benchmarks with a frozen Qwen3.5-9B, LiteEvo lifts pass@2 by 10.5 to 67.7pp and reaches comparable or higher pass@2 than a reproduction of HarnessX (71.0 against 67.3 on average) at 13.0 lower mean API cost. Harnesses evolved on train tasks keep their gains on unseen test tasks of four benchmarks, and LiteEvo also lifts Claude Code with Sonnet 4.6 by 1.2 to 71.4pp.