🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Agents

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

First page
ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
The curator’s take

Yupeng Su, Jiayi Tian and Zheng Zhang (UC Santa Barbara) with Souvik Kundu (Intel) present ReFold, a training-free rendering layer that compresses what a long-horizon agent sees while keeping the full interaction history recoverable.

Ask this paper

Key points
01

Problem. Agents re-send an append-only history at every step, so cost grows until sessions overflow the context window. Predictive context managers add model calls, invalidate prefix caches and discard content permanently.

02

Two operators. Content an earlier turn already displayed is replaced by a stub, and turns the agent reports as finished are folded into a one-line note. Neither needs an auxiliary predictor.

03

Cache-aware and reversible. Chunked rendering rewrites the cached prefix only every few steps, and every removal can be restored from the stored history at the cost of one restore.

04

Results. Across five long-horizon benchmarks and two frontier LLMs, ReFold cuts token use by up to 2.5x and halves KV-cache memory per session without lowering task success, and avoids up to 92% of forced compactions under capped budgets.

05

Serving. Under concurrent workloads it removes up to 100% of request queuing delay, speeds inference by up to 1.7x and cuts inference cost by up to 3.4x.

Abstract

Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.

Every Monday
Get next week’s papers.
Subscribe on Substack