🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Efficiency · Agents · Code

Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents

First page
Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
The curator’s take

Ritul Satish, Prasoon Sinha, Akiho Kawada and Neeraja Yadwadkar at UT Austin run nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench to separate the three decisions in a context compression policy (mechanism, trigger, and amount removed) and measure how each affects success, tokens, latency and cost.

Ask this paper

Key points
01

Motivation. In a trace of 13 million GitHub Copilot sessions, sessions that need compaction account for 44.2% of tokens served, and the median compaction removes 72.8% of context.

02

Fewer tokens can mean slower runs. On Terminal-Bench with Qwen, policies using about a third of the tokens can take 20% to 80% longer than running with full context, because summarization calls and extra steps add time.

03

Trigger choice matters. Step-triggered policies cut tokens per step the most but need 10% to 27% more calls. Stacked threshold policies cut tokens by 22% to 55% with call counts near full context.

04

Results do not transfer across models. The same policy (OTRC) reaches 51.3% with Qwen but drops Devstral to 38.7% and raises its latency. Raising the trigger threshold from 10K to 20K lifts Qwen from 50.7% to 73.3% on a SWE-bench subset while Devstral stays near 72%.

05

Takeaway. Compression policies need to be tuned per model and evaluated on latency and cost, not token count alone.

Abstract

As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to compress, and how much to remove into fixed policies. A systematic characterization is needed to disentangle these decisions and reveal how each affects task success and execution cost. We systematically vary these decisions across three open-weight models on SWE-bench Verified and Terminal-Bench 1.0. Across nearly 35,000 agent runs, we measure task success, token use, end-to-end latency, and estimated cost. We find that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent. Policies with similar overall success can solve different tasks, while the same policy can perform quite differently across models. Our results motivate evaluating compression by its effects on agent execution and tailoring policies to the task, model, and workload.

Every Monday
Get next week’s papers.
Subscribe on Substack