Trained Agentic Context Management

Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.
Ask this paper
Minimal harness. The only tools are self-invocation and range reads, chosen so the harness enables delegation without prescribing a decomposition strategy. Training data is synthetic oracle SFT generated by Claude Opus and Fable models.
Main result. With an 8,000-token per-agent context, the fine-tuned model matches GPT-5.4 (medium reasoning, 1M context) on OOLONG-synth once documents exceed 40K tokens, the length at which single-context performance starts to fall.
Untrained models misuse the harness. Base Qwen, GPT-5.4 and Kimi K2.6 delegate only sporadically or forward the original question unchanged. Opus 4.6 recognizes context pressure but builds subagent trees large enough to spend $300 on a 100-question eval.
Limits. On RULER the fine-tune stays above 85% but trails GPT-5.4, mainly because it chooses a lossy aggregation strategy on the common-words task. The author argues the recipe belongs inside frontier model training, with RL as the next step.
Abstract
We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a diverse synthetic dataset using this harness. With only 8,000 tokens of context, our small model is as strong as GPT-5.4 with 1M tokens of context on the OOLONG-synth benchmark when document length exceeds 40K tokens.