🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Agents · Safety

When History Fails to Become Experience: Action Calibration in Language Agents

First page
When History Fails to Become Experience: Action Calibration in Language Agents
The curator’s take

Jingyu Liu and Yong Liu (Renmin University of China) with Zhiwen Wang, Yuxin Jing and Huanyu Zhou (ByteDance) study how language agents use their own interaction history and find that they rarely connect each action to its outcome.

Ask this paper

Key points
01

History helps for the wrong reason. History improves task completion, but much of the benefit remains when past actions are shuffled, and adding reference history lowers recovery from failed steps from 15.8% to 11.4%.

02

Outcome labeling. Marking each observation as the outcome of the action that produced it, with no new information added, improves task success and reduces repeated actions.

03

Self-written experience hurts. Asking the model to summarize lessons from its own history lowers success relative to plain history, because some lessons restate feedback and others add unsupported judgments.

04

Learned calibrator. A calibrator trained on how each experience changes next-action quality across several actors decides when to write a lesson and what it says. It raises held-out success from 47.21% to 52.21% averaged across environments and seeds.

Abstract

Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we examine how agents use history. We find that history improves task completion overall, yet much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past actions and selectively records experience to guide subsequent decisions, improving task success beyond outcome labeling alone.

Every Monday
Get next week’s papers.
Subscribe on Substack