🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Memory · Multimodal · Agents

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

First page
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
The curator’s take

Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

Ask this paper

Key points
01

Gap. Existing multimodal memory benchmarks use personal-life photos and recall questions, often with fewer than three evidence images per question, and frontier models already exceed 90% on LoCoMo.

02

Benchmark. 1,000 expert-reviewed questions across Current State, Past State, Derived State, Change History and Conflict/Refusal, over information-dense artifacts that are revised many times. A generation harness separates state-transition synthesis from conversation filling.

03

Results. Across 27 configurations of frontier and open-weight models and memory methods, the best baseline (Gemini 3.7 Flash with hints) scores 43.4%; it reaches 61.7% on Current State but 27.8% on Change History.

04

What makes it hard. Removing statefulness raises matched-category accuracy by 43.2 points, against 12.6 for lower density and 8.2 for converting visuals to text; 89.7% of failures are state-reconstruction errors such as lost updates and stale values.

05

Mitigations. More reasoning effort and generic memory-management methods give limited gains, while state-aware designs do better.

Abstract

Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.

Every Monday
Get next week’s papers.
Subscribe on Substack