🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Memory · Evaluation

The Fragility of Self-Improving Agents

First page
The Fragility of Self-Improving Agents
Paper summary

Memory-based self-improving agents report gains that have never been checked against evaluation noise. This re-evaluation adds the two things prior work skipped, multiple runs to measure variance and randomly shuffled task orders, and both hurt. Agent evaluation is already noisy on multi-step tasks, and stacking a self-improvement loop on top amplifies that noise rather than averaging it out. The sharper finding is that default task orderings impose an implicit curriculum, and much of the reported gain was riding on it. Adding detailed rubrics and environment feedback to memory construction recovers part of the drop, and a significant gap remains. If you are measuring your own memory loop, shuffle the task order first.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack