🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 2, 2026
Agents

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

First page
Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
The curator’s take

Jeffrey Willette, Krishna C. Puvvada and Boris Ginsburg at NVIDIA introduce Long-Transduction, a controlled diagnostic for whether a model can keep applying state-dependent operations correctly across a long generation, the basic capability long-horizon agents depend on.

Ask this paper

Key points
01

Task design. The model must read, update and output input-dependent results (arithmetic, sorting, variable lookups, table transformations) over thousands of outputs, like an agent reconciling a long ledger.

02

Three independent axes. Context length, input formatting and local task complexity are varied separately, so each failure can be attributed to one cause.

03

Results on seven open-weight models. Accuracy drops 62.8% when context grows from 4K to 128K, 36.5% when only the input format changes, and 39.9% when local task complexity increases.

04

Implication. A model can accept a long input and still lose its place or stop applying an operation consistently partway through generation, so long-context acceptance does not establish long-horizon reliability.

Abstract

Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.

Every Monday
Get next week’s papers.
Subscribe on Substack