FlowState: Execution State as Memory for Long-Horizon LLM Agents

Minghao Li, Bangyan Li and colleagues at Ant International, Ant Group propose FlowState, an agent memory framework that stores execution state as typed, linked nodes with references back to raw tool output, so an agent can revisit earlier decisions and their evidence without carrying the full history in context.
Ask this paper
Representation. Requirements, environment facts, agent judgements and artifacts become typed state nodes with explicit relations and pointers to the original observations. Retention is separated from what the agent currently sees.
Two operations. Incremental State Update maintains the current state from new inputs and feedback; Progressive State Access reveals older states and their linked evidence step by step as the task needs them.
Results. With DeepSeek-V4-Flash, FlowState beats the full-context baseline by 4.55 points average success on MemoryArena and 13.95 points pass rate on tau-Bench, while using 43.2% and 40.6% fewer tokens.
Per-domain gains. On tau-Bench Airline the pass rate rises by 20.0 points with 61.2% fewer tokens. On Bundled Web Shopping it scores 42.89%, 11.78 points above the best memory-system baseline.
Across models. All three tested backbones get higher mean success and progress with 28.3% to 56.3% lower relative token use.
Abstract
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.