🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 27, 2026
Reasoning · Agents · Efficiency

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

First page
When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression
The curator’s take

Mingxuan Wang and colleagues (TierFlow team, Renmin University Gaoling School) study when an agent can safely drop its earlier reasoning, and propose ICLR (Interaction Aware Compression for Long Horizon Reasoning), a training-free online method.

Ask this paper

Key points
01

Method. Ranks past reasoning blocks by the entropy of a frozen proxy model and removes low-value ones, while always keeping actions, tool calls and observations.

02

Results. On 260 WorkBuddyBench tasks, average reward rises from 0.699 to 0.718 while input, output and cache-read tokens fall by 25.5%, 14.4% and 33.3%.

03

Trajectory amplification. Deleting reasoning locally changes later actions, so total compute changes nonlinearly; compression for agents cannot be evaluated one step at a time as in static chain-of-thought compression.

04

When forgetting is safe. Probing, activation patching and controlled trajectories indicate reasoning becomes replaceable once the derived state has been written out to code, files, tool outputs or environment feedback.

Abstract

Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 to 0.718, while reducing input, output, and cache read tokens by 25.5%, 14.4%, and 33.3%, respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history.

Every Monday
Get next week’s papers.
Subscribe on Substack