🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Agents

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

First page
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
The curator’s take

Yuyao Ge and colleagues at the Institute of Computing Technology, Chinese Academy of Sciences (with UC Merced and Tsinghua) present SkillForge, an agentic RL method in which the skill library and the policy are updated together. Accepted at NeurIPS 2026.

Ask this paper

Key points
01

Problem. Append-only skill libraries accumulate obsolete or harmful entries. The authors name delayed obsolescence: a skill that once reached the stable state but loses fitness as the policy improves.

02

Lifecycle. Each skill moves through trial, active, stable and retired states based on measured fitness. Before RL, the base model's own rollouts pre-retire low-fitness skills, and the filtered library seeds supervised fine-tuning.

03

Co-evolution. During RL, each iteration applies selective retirement, stabilization and LLM-guided mutation to the library alongside policy optimization.

04

Results. 92.4% on ALFWorld and 78.4% on WebShop, 2.8% and 7.8% relative over SkillRL, with the library kept compact; gains also hold on search-augmented QA.

05

SkillFurnace. A released dataset of 5,852 records: retirement-filtered SFT trajectories, library snapshots with fitness, and 318 retirement events with human-annotated failure categories.

Abstract

Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.

Every Monday
Get next week’s papers.
Subscribe on Substack