🚀NEW LABGetting Started with Claude AgentsStart lab
Code

Self-Taught Optimizer (STOP)

First page
Self-Taught Optimizer (STOP)
Paper summary

Proposes recursively self-improving code generation where an LLM-scaffolded program improves itself.

Ask this paper

Key points
01

Seed improver: A "seed improver" program first improves an input program to return the best solution found - a self-improvement scaffold built on GPT-4.

02

Recursive improvement: The seed improver is itself tasked with improving itself, producing the first concrete demonstration of recursive self-improvement in LLM code generation.

03

GPT-4 capable: Shows that GPT-4 models can write code that modifies itself iteratively, producing measurably better scaffolds than the initial seed.

04

Foundational work: An early, influential demonstration of the LLM-as-code-modifier pattern that would reappear across 2024 in agent and tool-use research.

Every Monday
Get next week’s papers.
Subscribe on Substack