Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

Jeffrey Willette, Krishna C. Puvvada and Boris Ginsburg at NVIDIA introduce Long-Transduction, a controlled diagnostic for whether a model can keep applying state-dependent operations correctly across a long generation, the basic capability long-horizon agents depend on.
Ask this paper
Task design. The model must read, update and output input-dependent results (arithmetic, sorting, variable lookups, table transformations) over thousands of outputs, like an agent reconciling a long ledger.
Three independent axes. Context length, input formatting and local task complexity are varied separately, so each failure can be attributed to one cause.
Results on seven open-weight models. Accuracy drops 62.8% when context grows from 4K to 128K, 36.5% when only the input format changes, and 39.9% when local task complexity increases.
Implication. A model can accept a long input and still lose its place or stop applying an operation consistently partway through generation, so long-context acceptance does not establish long-horizon reliability.
Abstract
Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.