šŸš€NEW LABGetting Started with Claude AgentsStart lab
← All papers Ā /Ā  Oct 2, 2026
Agents

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

First page
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
The curator’s take

Prithwish Jana (Georgia Tech) with Nikos Kanakaris, Sahika Genc and colleagues at AWS AI Labs, plus CMU and WashU, introduce MILO (Meta-evolutionary Island Orchestration), an automated harness-discovery framework that evolves the search strategy along with the harness it is searching for.

Ask this paper

Key points
01

Island lineage memory. Each island grows from a different seed harness and keeps rejected candidates as negative evidence, so failed directions are not retried and the search does not collapse onto a few designs.

02

Whole-harness mutators plus an orchestrator. Per-island mutator agents rewrite the complete harness from parent failure traces; when an island stalls, an orchestrator agent grafts or splits lineages, reassigns mutators and revises the task curriculum.

03

Results with Opus 4.8. Gains over the initial harness of +12.0%, +28.3% and +10.3% on Terminal-Bench 2.1, PaperBench and DeepSWE, against best prior-search gains of +4.5%, +18.3% and 0%. On Terminal-Bench 2.1 it reaches 86.1 ± 2.0%, above the leaderboard's top entry (83.8 ± 2.3%), with 26% fewer tokens than its starting harness.

04

Multi-objective fitness. Candidates are admitted on Pareto gain over accuracy, tokens and latency, which most evolutionary harness searches ignore.

05

Transfer and math. Evolved harnesses transfer to Frontier-Bench without further search, and the same machinery tightens three best-known bounds on EinsteinArena (Erdős minimum overlap and two autocorrelation inequalities).

Abstract

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).

Every Monday
Get next week’s papers.
Subscribe on Substack