🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Memory · Agents

AMBER: Training Long-Horizon Web Agents through Append-Only Memory

First page
AMBER: Training Long-Horizon Web Agents through Append-Only Memory
The curator’s take

Chinmay Savadikar, Tianfu Wu, Lingyun Wang and colleagues at North Carolina State University and Shopify introduce AMBER, an append-only memory that a web agent learns to write while it reasons and acts, trained end-to-end with outcome-reward RL.

Ask this paper

Key points
01

Problem. Trained overwrite memories must carry each fact through every later rewrite, which is hard to learn from sparse rewards; the authors find they delete key facts and corrective feedback from the environment.

02

Method. The agent jointly learns to reason, act and write free-form memory entries, and an append-only rule guarantees nothing written is lost, with no need for large curated SFT data.

03

Results. On WebArena Lite AMBER improves average success over overwrite memory (MemAgent) by 4.09 points and raises the fraction of tasks solved in all five repeated runs by 4.8 points; with Qwen 3.5 27B it beats Haiku 4.5 by 2.1 points.

04

Supervision cost. AMBER matches an overwrite baseline trained on substantially more expensive curated supervision.

05

Behaviour. After an environment correction, MemAgent re-issues the invalid action six or more times in 15.4% of events, against 3.3% for AMBER.

Abstract

Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.

Every Monday
Get next week’s papers.
Subscribe on Substack