Self-Taught Optimizer (STOP)
Free while signed in. Answers cite the passages they came from.

Proposes recursively self-improving code generation where an LLM-scaffolded program improves itself.
Seed improver: A "seed improver" program first improves an input program to return the best solution found - a self-improvement scaffold built on GPT-4.
Recursive improvement: The seed improver is itself tasked with improving itself, producing the first concrete demonstration of recursive self-improvement in LLM code generation.
GPT-4 capable: Shows that GPT-4 models can write code that modifies itself iteratively, producing measurably better scaffolds than the initial seed.
Foundational work: An early, influential demonstration of the LLM-as-code-modifier pattern that would reappear across 2024 in agent and tool-use research.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack