🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Agents · Evaluation

When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds

First page
When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
The curator’s take

Ke Wang (Cambridge, Georgia Tech), Zijie Zhao (MIT) and colleagues formalize Improvement Fidelity, the requirement that an update a proxy verifier scores as better is still better after it is deployed into a world that reacts to it, and propose PIVOT-KG to spend a small budget of high-fidelity evaluations where they can change the update decision.

Ask this paper

Key points
01

Problem. Self-improving agents propose policy updates and score them with a proxy. When deployment changes the environment, such as other agents responding, an update that looks better to the proxy can perform worse after release.

02

Accuracy is not enough. The authors show that a verifier can rank policies accurately overall and still choose the wrong replacement, because errors on the specific proposed updates and small candidate margins decide the outcome.

03

Measured reversals. In Leduc poker, the fraction of proxy-positive updates that become negative after deployment rises from 2.5% with one-step opponent response to 51.7% with eight-step adaptation.

04

PIVOT-KG. A paired validator that allocates scarce high-fidelity evaluations by expected reduction in selection regret, instead of spreading them uniformly. On long-response Leduc it lowers regret from 0.046 to 0.009 with two queries; gains on Kuhn and Melting Pot are small or not significant.

Abstract

Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. We formalize this gap as Improvement Fidelity, which asks whether proxy improvements preserve the sign and ordering of deployment improvements over the updates an improvement process actually proposes. We show that global policy accuracy need not guarantee update fidelity: operator shift and deployment response can create update-level errors, while candidate margins determine whether those errors change the replacement decision. We introduce PIVOT-KG, a paired, decision-aware validator that allocates scarce high-fidelity evaluation according to the expected reduction in selection regret per unit cost. Across 90 held-out roots in Leduc, Kuhn, and Melting Pot, proxy and deployment optimal sets are disjoint in 51 cases. In an eight-candidate HighwayEnv stress test, PIVOT-KG reduces mean improvement-selection regret from 0.0435 under the exact Uniform validation rule to 0.0055 at the primary budget. Together, these results show why reliable self-improvement should evaluate proposed improvements in the worlds they induce, while providing a practical rule for allocating scarce deployment evidence when it can affect the replacement decision.

Every Monday
Get next week’s papers.
Subscribe on Substack