🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Agents · Efficiency · Evaluation

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

First page
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
The curator’s take

Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan, Shuowei Jin and Beidi Chen (Carnegie Mellon, Infini-AI Lab) introduce ServeLearnBench, a benchmark for agents that must infer and revise hidden environment policies from serving experience.

Ask this paper

Key points
01

Setting. An evolving-environment streaming dataset where hidden policies change over time and agents only get interaction and outcome feedback. It covers retail support, banking and sales-pitch generation, with 53 environment windows and 7,718 tasks.

02

Scale of evaluation. Five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness and Prime) across six models, for 28 model-harness pairs and 252 learning runs.

03

Capability vs learning. With the policy disclosed, models reach 95.4% on hidden cases on average; learning it from experience, the best pair reaches 59.6% on Retail L3 and the median pair 22.0%.

04

Harness differences. Prime and Continual Harness adapt fastest, while RAG, Mem0 and SkillOpt trail by 13 to 20 points. Continual Harness reaches 57.7 test reward at 52% of Prime's test cost.

05

Costs and regressions. Adaptation can degrade behaviour that was already correct, and insufficient exploration is the main bottleneck the authors identify.

Abstract

Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.

Every Monday
Get next week’s papers.
Subscribe on Substack