ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Sungho Park (POSTECH, intern at Microsoft), Jue Zhang, Pengfei Gao and colleagues at Microsoft, POSTECH and KAIST introduce ActiveSaddler, which adapts the training scenarios used to drive automated harness optimization as the harness changes.
Ask this paper
Gap. Harness optimizers study how to update the harness but fix which scenarios generate feedback, even though the useful scenarios shift as the harness improves.
Non-stationary bandit. Recurring failures are abstracted into failure-pattern arms; the method estimates learning progress per pattern and balances revisiting known weaknesses against exploring new scenarios.
Results. On GAIA2 and Terminal-Bench 2.0, test Pass@1 improves by 4.4 and 7.5 points over the same optimizer with a fixed scenario order.
Ablations. Gains depend on building targets dynamically, tracking their changing utility, and keeping some budget for discovering new failures.
Abstract
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.