Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Hongqiang Lin (Zhejiang University) with Chao Liu, Xipeng Cao and colleagues at Alibaba Group introduce EvoPathBench, a benchmark that measures self-evolving agents at each checkpoint of their memory or skill updates instead of only at the end.
Ask this paper
Protocol. The base model and tools are fixed, the evolving artifacts are frozen at successive checkpoints, and each target capability is tested on held-out episodes built from public trading data and calibrated trajectories.
Three capabilities. Generalization to unseen tasks, retention after unrelated learning, and adaptation of rules to new evidence.
Findings. Gains on similar unseen tasks often weaken under distribution shift; retention losses concentrate in a minority of evolution paths; no evaluated method achieves reliable rule adaptation.
Selection is the bottleneck. Self-evolution generates candidate artifacts with large held-out gains, but the updates the agents actually select fall short of that potential, which points to candidate evaluation and selection as the component to improve.
Cost trade-off. Among skill-evolution methods, SkillBoost achieves the highest gain but needs the token budget of SkillOpt for a 45.5% larger gain.
Abstract
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose \textsc{EvoPathBench}, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.