🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 29, 2026
Agents · Code · Training

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

First page
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
The curator’s take

Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

Ask this paper

Key points
01

LOHA layout. Latent Observations, Hard Actions: older tool outputs become soft tokens, while the agent's turns and the last K observations stay in plain text for exact reference.

02

Anchored Context Distillation. Trains the model to read the latent view by distilling full-text predictions into it, while anchoring behavior on plain-text inputs to the base model to limit drift.

03

Compression and cost. With K=3 on SWE-bench Verified, context per call drops 43% for Qwen3-4B and 57% for SWE-Master-4B-RL; resolve rates fall from 14.5% to 12.1% and 27.5% to 21.8%.

04

Recency window. At K=8, resolve rates recover to 14.4% and 23.0%; larger windows trade compression for task success.

05

Under a context limit. With a 32K-token cap, the compressed agent resolves 21.1% of a 199-instance subset versus 11.1% for full text, at 1.9x the serving throughput.

Abstract

Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text needed for exact reference, with Anchored Context Distillation (ACD), a training method that enables latent reading while constraining behavioral drift. LOHA compresses older tool observations into soft tokens while retaining the agent's own turns and the last K observations in text, providing compact access to historical information and exact access to recent content. To enable the agent to use this representation, ACD distills the base model's full-text predictions into the latent view while anchoring its behavior on plain-text inputs to the same base model. On SWE-bench Verified, K=3 reduces context per call by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. A single-run recency sweep reaches 14.4% and 23.0% at K=8, with larger windows generally favoring task performance over compression. Under a 32K-token limit, Qwen3 with K=3 resolves 21.1% of a 199-instance subset versus 11.1% for the same adapted agent using full text. In concurrent single-GPU serving, it achieves 1.9 times that full-text agent's instance throughput.

Every Monday
Get next week’s papers.
Subscribe on Substack