🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Efficiency · Agents

FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

First page
FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
The curator’s take

Shantanu Dixit, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan and colleagues at M365 Research, Microsoft introduce FOCUS, a training-free test-time method that compresses an agent's history by keeping the interaction units that causally shape its next decisions.

Ask this paper

Key points
01

Reframed objective. Instead of learning offline what to discard (contrastive guidelines, distilled compressors, trained policies), FOCUS asks which past interactions the agent's future decisions depend on, and treats compression as preserving those decisions over discrete interaction units.

02

No training. It needs no offline data or fine-tuning and attaches to closed-API frontier models as a separate compression layer.

03

Results. Across API and tool-calling, QA, web and multi-turn dialogue benchmarks it cuts peak context by up to 48% and dependency by 73% while raising task success by up to 8.9 points over uncompressed execution.

Abstract

LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.

Every Monday
Get next week’s papers.
Subscribe on Substack