SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

Guanyu Nie, Mingxuan Yuan and colleagues at Huawei Noah's Ark Lab treat repeated agent skill updates as a training process that can overfit, and introduce SkillEvoReg to regularize it.
Ask this paper
Skill-evolution overfitting. Repeated edits to external skills accumulate redundant or task-specific rules, and new edits can break behavior that previously worked.
Three regularizers. Skill dropout during update generation, complexity-aware local regularization to limit structural growth, and causal counterexample validation that checks each candidate update for regressions.
SkillOpt results. On SpreadsheetBench, the final test score is 59.25 versus 47.00 for SkillOpt, with the skill shrinking from 6,060 to 2,341 tokens; LiveMath falls from 36.41 to 30.18, a drop the authors trace to one answer pattern.
Transfer. On SkillEvolBench composition tasks, success reaches 26.67% versus 18.33% after one unregularized pass and 13.33% after two.
Plug-in design. It wraps each system's own skill evolver and evaluator, and flags regressions that size metrics alone miss.
Abstract
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neural-network training. SkillEvoReg combines training-time skill dropout, which perturbs update generation, and complexity-aware local regularization, which controls unnecessary structural growth, with causal counterexample validation (CCV), which provides targeted behavioral validation of candidate-specific regressions. We instantiate the framework across heterogeneous skill-evolution systems while retaining each system's native skill evolver and task evaluator. Across SkillOpt, SkillEvolBench, and ContinualSkillBench, SkillEvoReg consistently controls skill-state growth while preserving competitive downstream capability, improves several transfer and later-stage evolution outcomes, and identifies update-level regressions that structural metrics alone cannot reveal. These results suggest that explicit regularization is a useful complement to increasingly capable skill updaters.