🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 2, 2026
Agents

Learning from Research: Toward Lifelong Agent Harness Evolution

First page
Learning from Research: Toward Lifelong Agent Harness Evolution
The curator’s take

Jingbo Yang (UCSB, intern at Microsoft), Kwei-Herng Lai, Evgeniy Gabrilovich, Shiyu Chang and colleagues at Microsoft and UC Santa Barbara introduce ScholarEvolve, which uses published agent research as the source of candidate changes when evolving a harness around a fixed model.

Ask this paper

Key points
01

Problem with reactive evolution. Existing harness optimizers let a meta coding agent edit the harness from observed failures, which limits changes to what that agent already knows.

02

Literature as search space. ScholarEvolve splits the harness into functional modules, runs topic modeling over recent papers to find distinct improvement strategies per module, implements them and evaluates combinations.

03

Lifelong design. New publications can be added over time, so research advances become candidate harness updates without waiting for a failure to motivate them.

04

Results. Qwen3.5-27B goal completion on AppWorld Challenge rises from 49.6% to 63.6%, and GPT-5.4-mini pass@1 on Tau2-Bench Telecom rises from 72.7% to 81.9%.

Abstract

Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.

Every Monday
Get next week’s papers.
Subscribe on Substack