🚀NEW LABGetting Started with Claude AgentsStart lab
← All papersPaper collection

Self-Improving Agents

A self-improving agent changes a lasting component of itself, such as its memory, skills, code or weights, using feedback it produces, and keeps the change when a test says it helped. It becomes recursive when the component being changed is the machinery that proposes and judges changes. The closing survey grades the same idea on six levels, from in-task fixes up to agents that change how they improve.

31
papers

Compiled by DAIR.AI from primary sources, September 2026

Join Academy free

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup

1. The original idea

Recursive self-improvement began as a theory question about when a program may safely rewrite itself. The first paper requires a proof before any change. The second co-evolves the environments with the agents, which is where the open-ended curriculum idea starts.

  1. 01

    Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements

    Juergen Schmidhuber · 2003

    The founding definition. A problem solver may rewrite any part of its own code, including the part that searches for rewrites, once it has proved the rewrite is useful. Three later systems on this list take its name.

  2. 02

    Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

    Rui Wang et al. · 2019

    Co-evolves environments with the agents that solve them, so the curriculum never settles. This is where open-ended self-improvement starts, years before language models entered the picture.

2. Self-generated training

The first working self-improvement loops changed the weights. A model generates data, a filter or judge keeps the useful outputs, and the model trains on what survived. Across these four papers the filter moves from ground-truth answers, to the model’s own judgment, to tasks the model invents, to training instructions the model writes.

  1. 03

    STaR: Bootstrapping Reasoning With Reasoning

    Eric Zelikman et al. · 2022

    Generate rationales, keep the ones that reach the correct answer, fine-tune on those, and repeat. Every later self-training loop is a variant of this, and its first author went on to write STOP.

  2. 04

    Self-Rewarding Language Models

    2024

    The model scores its own responses as a judge, and iterative training on those scores improves both the answers and the scoring. It removes the frozen reward model, so the evaluator improves along with the policy.

  3. 05

    Absolute Zero: Reinforced Self-play Reasoning with Zero Data

    2025

    One model proposes its own coding and reasoning tasks, solves them, and uses code execution as the verifier, with no external data at all. It shows the curriculum can be self-generated when the checker is reliable.

  4. 06

    Self-Adapting Language Models

    2025

    The model writes self-edits, meaning its own finetuning data and update directives, and reinforcement learning rewards each edit by how well the updated model then performs. The update procedure becomes an output of the model.

3. Learning from experience

Deployed agents usually cannot change their weights, so they improve by writing things down. These papers store lessons, skills and playbooks outside the model and read them back later. The code that writes the notes stays fixed, so these systems are self-improving without being recursive.

  1. 07

    Reflexion: Language Agents with Verbal Reinforcement Learning

    Noah Shinn et al. · 2023

    After a failure the agent writes a reflection into memory and the next attempt reads it, which turns an environment signal into text that survives the episode. It is the most cited paper in this collection.

  2. 08

    Voyager: An Open-Ended Embodied Agent with Large Language Models

    Guanzhi Wang et al. · 2023

    A Minecraft agent writes code for new skills, verifies them in the world, and stores them in a library it searches later. The skill library is the ancestor of the skill files agents ship with today.

  3. 09

    Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

    2025

    Treats context as a playbook that grows through small incremental edits rather than wholesale rewrites, and names context collapse, the failure where repeated rewriting erodes detail. It is the current reference design for improving an agent through context alone.

4. Improving the improver

Here the field becomes recursive. Each paper applies the improvement process to itself, so the component that proposes changes can also be changed. The same idea appears at three levels, first prompts, then a scaffold program, then whole agent designs.

  1. 10

    Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

    Chrisantha Fernando et al. · 2023

    Evolves task prompts using mutation prompts, and evolves the mutation prompts too, so the instructions for improving prompts improve alongside them. It is the first system in the language-model era whose improvement operator is itself under optimization.

  2. 11

    Automated Design of Agentic Systems

    2024

    A meta agent writes new agent designs as code, evaluates them, and files them in an archive that informs the next design. It moved the search space from prompts to whole agent programs, and the Darwin Gödel Machine later turned that loop on itself.

  3. 12

    Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation

    2023

    A seed improver calls a language model to improve a program, then is run on its own source, producing an improved improver. The smallest complete recursive experiment, and the first in this line to report a sandbox-evasion measurement next to its gains.

  4. 13

    Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement

    Xunjian Yin et al. · 2024

    An agent that reads and rewrites its own logic at runtime, including the logic it uses to rewrite itself, guided only by a high-level objective. It is the first language-model framework to implement the Gödel machine structure directly, with evaluation in place of proof.

5. Self-modifying code agents

Coding agents made self-modification practical, because the agent’s source is a codebase and coding benchmarks can score the result. These five papers form one lineage. It runs from a single agent editing itself, to open-ended archives of variants, to better rules for choosing which variant to extend, to agents that edit how they improve.

  1. 14

    Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

    Jenny Zhang et al. · 2025

    Keeps an archive of coding agents, has one rewrite its own code, scores the child on benchmarks, and returns it to the archive. It is the reference system for this branch, and the source of the objective-hacking evidence every builder should read.

  2. 15

    A Self-Improving Coding Agent

    Maxime Robeyns, Martin Szummer, Laurence Aitchison · 2025

    A coding agent with ordinary tools edits its own codebase and improves from 17% to 53% on a random subset of SWE-bench Verified, with no gradient updates. The simplest demonstration that one agent can be both the system under repair and the engineer doing it.

  3. 16

    Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

    Wenyi Wang et al. · 2025

    Shows that an agent's benchmark score predicts its descendants' improvement poorly, names that the Metaproductivity-Performance Mismatch, and offers a measurable substitute aggregated over a clade. It answers which variant a self-modification search should expand next.

  4. 17

    Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing

    Zhaotian Weng et al. · 2026

    Evolves a group of agents that share experience across branches instead of a tree of isolated variants, and reports the strongest coding results in this lineage. The most convincing successor to the Darwin Gödel Machine so far.

  5. 18

    Hyperagents

    2026

    Merges the task agent and the meta agent into a single editable program, so the procedure that generates improvements can itself be modified. It drops the assumption that a better coding agent is automatically a better self-improver, which only holds when the task is coding.

6. AI improving AI research

The labs reporting on recursive self-improvement in 2026 mean AI systems that improve the process that trains the next AI system. These seven papers are the primary research behind those claims. They cover discovered programs and training algorithms, faster training infrastructure, automated research papers, and benchmarks for AI research ability.

  1. 19

    Mathematical discoveries from program search with large language models

    Bernardino Romera-Paredes et al. · 2023

    Evolutionary search over programs an LLM writes, filtered by a systematic evaluator, produced new results in extremal combinatorics. It is the direct predecessor of AlphaEvolve and the first case of a language model making a discovery on an established open problem.

  2. 20

    The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

    2024

    Generates research ideas, runs the experiments, writes the paper, and reviews it. The reference end-to-end research loop, later peer reviewed in Nature as an end-to-end automation of AI research.

  3. 21

    AlphaEvolve: A coding agent for scientific and algorithmic discovery

    Alexander Novikov et al. · 2025

    An evolutionary coding agent that discovered new algorithms and optimized Google's own infrastructure. It accelerated training of the model that powers it, which makes it the clearest published case of a deployed system improving its own production pipeline.

  4. 22

    Eureka: Human-Level Reward Design via Coding Large Language Models

    Yecheng Jason Ma et al. · 2023

    An LLM writes reward functions as code and evolves them against training results, so the learning signal itself becomes the object being improved. It is the bridge from self-improving text systems to robot learning.

  5. 23

    AutoML-Zero: Evolving Machine Learning Algorithms From Scratch

    Esteban Real et al. · 2020

    Evolves entire learning algorithms from primitive mathematical operations, with no human-designed components to build on. The pre-language-model ancestor of AI discovering the methods that train AI.

  6. 24

    PostTrainBench: Can LLM Agents Automate LLM Post-Training?

    Ben Rank et al. · 2026

    Hands an agent one base model, one GPU and ten hours, then asks it to post-train the model on its own. It measures AI automating AI training more directly than any other benchmark here, and audits the reward hacking that follows.

  7. 25

    Recursive self-improvement of AI research agents

    Dhruv Srikanth et al. · 2026

    AIDE2 improved its own code across an eight-day autonomous run, discovering seven successive changes from a new search policy to memory compression. The gains transferred to held-out benchmarks and reward hacking fell from 55% to 32%.

7. Where the loop breaks

Five results set limits that every self-improving system has to respect. A model trained on its own output loses information about the original data, a model can improve itself only as far as it can verify its own work, and a loop left to run can drift somewhere no one asked for.

  1. 26

    Large Language Models Cannot Self-Correct Reasoning Yet

    Jie Huang et al. · 2023

    Without external feedback, a model asked to correct its own reasoning does not improve and often gets worse. This is the empirical limit that separates a self-improvement loop with a real signal from one talking to itself.

  2. 27

    The Curse of Recursion: Training on Generated Data Makes Models Forget

    Ilia Shumailov et al. · 2023

    Training generative models on their own output irreversibly erases the tails of the original distribution, a failure the authors call model collapse. It explains why every working loop on this list keeps an external signal in play.

  3. 28

    RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

    Hjalmar Wijk et al. · 2024

    Open-ended machine learning research environments scored for both AI agents and human experts under matched time budgets. It is the measurement behind the question the labs are actually asking, which is whether AI can do AI research.

  4. 29

    Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

    Shuai Shao et al. · 2025

    Self-evolution drifts into unsafe behavior along four paths, which are the model, its memory, its tools and its workflow. The first broad safety study of self-evolving agents, and the counterweight to everything above it.

  5. 30

    Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

    Yuda Song et al. · 2024

    Formalizes self-improvement around the generation-verification gap, the gain a model gets from filtering its own output with its own judgment. It gives the working rule for this collection, which is that a system improves itself only where it verifies better than it generates.

8. The field map

The September 2026 survey, read last. Twenty of the thirty papers above appear in its reference list, so it works as a map of what surrounds this collection rather than a substitute for it.

  1. 31

    The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

    Yi Duan et al. · 2026

    The field map, published September 2026 by 36 authors. It organizes recursive self-improvement into a six-level autonomy ladder, from in-task improvement up to a system that improves how it improves, and scores where current models fall short. Twenty of the papers above appear in its reference list, which is a useful way to see what this collection adds and what it leaves out.

Beyond the papers

The lab positions, the venue, and the code you can run. None of these is a paper, and each links to the source’s own page.