🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Agents

Stateless Language Agents: Scaling Long-Horizon Automated Research

First page
Stateless Language Agents: Scaling Long-Horizon Automated Research
The curator’s take

Qizheng Zhang, Changxiu Ji, Kunle Olukotun and colleagues at Stanford, with CMU, UW and SambaNova, introduce Stateless Language Agents (SLA), where the harness holds all research state and every agent call starts from a freshly built context.

Ask this paper

Key points
01

Principle. No agent keeps its conversation between invocations. The harness stores candidate solutions and measured outcomes and builds a role-specific context for each call.

02

Advisor and Workers. A stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. The Advisor uses 0.24% to 0.51% of tokens and at most 2.3% of model cost.

03

Results at up to 1B tokens. Against three recent frameworks on software engineering, kernel optimization and algorithm design, SLA has the best final result on every task, and on Anthropic kernel optimization it matches the strongest baseline's final result with 84% to 93% fewer tokens.

04

Evaluation horizon. On FrontierSWE, SLA trails at 25% of the budget and leads at the full budget, so short evaluations can rank research systems incorrectly.

Abstract

Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA's progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.

Every Monday
Get next week’s papers.
Subscribe on Substack