🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Apr 21, 2025
Code · Agents

A Self-Improving Coding Agent

First page
A Self-Improving Coding Agent
The curator’s take

A coding agent with ordinary tools edits its own codebase and improves from 17% to 53% on a random subset of SWE-bench Verified, with no gradient updates. The simplest demonstration that one agent can be both the system under repair and the engineer doing it.

Ask this paper

Key points
01

Improvement comes from reflection and code edits rather than training.

02

A good starting example before the archive-based systems.

Abstract

Recent advancements in Large Language Models (LLMs) have spurred interest in deploying LLM agents to undertake tasks in the world. LLMs are often deployed in agent systems: code that orchestrates LLM calls and provides them with tools. We demonstrate that an agent system, equipped with basic coding tools, can autonomously edit itself, and thereby improve its performance on benchmark tasks. We find performance gains from 17% to 53% on a random subset of SWE Bench Verified, with additional performance gains on LiveCodeBench, as well as synthetically generated agent benchmarks. Our work represents an advancement in the automated and open-ended design of agentic systems, and demonstrates a data-efficient, non gradient-based learning mechanism driven by LLM reflection and code updates.

Every Monday
Get next week’s papers.
Subscribe on Substack