Self-Evolution of LLMs

This survey organizes the emerging literature on self-evolving LLMs - models that improve through their own generated experience rather than additional human supervision. The authors propose a unified four-phase cycle and taxonomize existing methods across both standalone models and agent systems.
Ask this paper
Four-phase cycle: Self-evolution is framed as iterative cycles of experience acquisition, experience refinement, model updating, and evaluation, mirroring how humans learn from practice.
Two application domains: The taxonomy separates standalone-LLM self-evolution (e.g., self-instruct, self-reward) from LLM-agent self-evolution (e.g., tool-use refinement, long-horizon planning).
Motivation: Reduce dependence on costly human annotation and break through plateaus as task complexity grows, with self-evolution positioned as a possible path toward more autonomous capability growth.
Open directions: The paper closes with concrete gaps - stable update dynamics, evaluation of self-evolved capabilities, and safety considerations as self-improvement loops tighten.