🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Training

Strategy Lock-In

Free while signed in. Answers cite the passages they came from.

First page
Strategy Lock-In
The curator’s take

Agents post-training other agents is one of the more load-bearing assumptions in current recursive self-improvement arguments. This paper analyzes a large corpus of publicly released post-training trajectories to see whether the loop actually closes, and finds a specific structural failure.

Key points
01

The first step decides everything: Across tasks, the agent locks in its training strategy at the very first step, then spends the entire remaining budget on local adjustments inside that choice. There is no mechanism for stepping back out.

02

Better scaffolding lifts execution, not strategy: An experience-driven scaffold was worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen throughout. The agent got better at the plan it had already committed to.

03

Human guidance does not survive training: Redirecting the opening choice by hand worked, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.

04

Why it matters: What agents lack here is a way to reconsider strategy while execution is still running. That is a different problem from reasoning quality or tool use, and none of the three escalating fixes tried here touch it, which sets a clear target for the next round of work.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack