🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 24, 2026
Agents · Memory · Evaluation

DolphinBench: Mapping the Pareto Frontier of Agent Memory

First page
DolphinBench: Mapping the Pareto Frontier of Agent Memory
The curator’s take

Soumil Rathi, Deshraj Yadav and Taranjeet Singh (Mem0) release DolphinBench, a memory benchmark that scores agents on actions they take in simulated apps rather than on answers to recall questions, and requires every submission to report cost and latency with accuracy.

Ask this paper

Key points
01

Action-based tests. Three knowledge-work personas each carry about 500k tokens of user messages; 200 tasks per persona require the agent to act (for example, edit a calendar event) using facts buried in that history, so the task does not announce what to retrieve.

02

Every task verified. Each test is kept only if an agent succeeds with the relevant history and fails without it.

03

Three metrics required. Accuracy, total cost across ingestion and testing, and median task latency are all mandatory, which lets the benchmark map a Pareto frontier.

04

Results. With the Hermes harness and Mem0, accuracy is 70.67% on GPT-5.6-Luna and 47.83% on MiniMax M3; with GPT-5.6-Luna, Mem0 beats built-in memory by 5 points and cuts median latency from 44.35 s to 37.69 s while cost rises from $61.48 to $96.21.

05

Rankings shift with the harness. Mem0 leads under Hermes and Honcho leads under Claude Code; note that the authors build Mem0.

Abstract

Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.

Every Monday
Get next week’s papers.
Subscribe on Substack