Self-Evolution of LLMs
Free while signed in. Answers cite the passages they came from.

This survey organizes the emerging literature on self-evolving LLMs - models that improve through their own generated experience rather than additional human supervision. The authors propose a unified four-phase cycle and taxonomize existing methods across both standalone models and agent systems.
Four-phase cycle: Self-evolution is framed as iterative cycles of experience acquisition, experience refinement, model updating, and evaluation, mirroring how humans learn from practice.
Two application domains: The taxonomy separates standalone-LLM self-evolution (e.g., self-instruct, self-reward) from LLM-agent self-evolution (e.g., tool-use refinement, long-horizon planning).
Motivation: Reduce dependence on costly human annotation and break through plateaus as task complexity grows, with self-evolution positioned as a possible path toward more autonomous capability growth.
Open directions: The paper closes with concrete gaps - stable update dynamics, evaluation of self-evolved capabilities, and safety considerations as self-improvement loops tighten.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack