Self-Improving Agents Survey
Free while signed in. Answers cite the passages they came from.

Self-improving agents are moving from research demos into deployed systems, and this survey gives the trend a clean formalism. It frames a modern agent as a foundation model coupled with an operational scaffold of prompts, memory, tools, and control logic, then treats self-improvement as a self-induced update that commits changes to either the weights or the scaffold.
Two update targets: Improvement splits into foundation-model updates to the weights and scaffolding updates to prompts, tools, memory, and control code, giving a shared vocabulary for work that usually looks unrelated.
Signals that drive change: The survey organizes methods by where the learning signal comes from, spanning intrinsic generative demonstrations, intrinsic evaluative feedback, and extrinsic exploratory experience in real or simulated environments.
Full-scaffolding frontier: The most open-ended methods rewrite the agent itself through self-referential code updates, generate-test-patch loops, and open-ended search over agent designs, pushing toward controllable evolution with little human input.
Why it matters: As teams wire agents to improve from their own experience, a single map of update targets, signals, and applications across software, web, gaming, science, and robotics turns a scattered literature into something builders can actually navigate.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack