🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 28, 2026
Memory

Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations

First page
Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations
The curator’s take

Zihao Lu, Zhihang Yuan and Lei Shi at Alibaba Cloud Computing introduce EGMemory, a memory system for long multi-party conversations that stores message-level evidence separately from an explicit, revisable state.

Ask this paper

Key points
01

State machine design. Persistent message evidence is kept apart from an active state. Writes go through a propose-verify-commit protocol so a state change is committed only when grounded in evidence, which handles information that participants later revise.

02

Read-time navigation. Queries iteratively resolve the relevant state and its supporting evidence, using conversational structure to narrow the search and lexical-semantic relevance to rank candidates.

03

No memory training. The system runs through prompting and tool use, with no memory-specific policy training.

04

Results. 68.2% on GroupMemBench and 77.9% on EverMemBench, 22.7 and 21.4 points above the strongest evaluated baselines. The strongest one-shot retrieve-then-read controls reach only 45.5% and 32.1%.

05

Two-party generalization. 73.6% on LoCoMo, 4.3 points above the strongest baseline.

Abstract

Long-horizon conversational memory is especially challenging in multi-actor settings, where relevant evidence is distributed across participants and contexts and previously established information may later be revised. We introduce EGMEMORY, which formulates long-horizon multi-actor memory as a searchable state machine that separates persistent message-level evidence from an explicit active state. At write time, adaptive state resolution and an evidence-grounded propose-verify-commit protocol govern how this state evolves. At read time, adaptive evidence navigation iteratively resolves the state and supporting evidence required for a query, using conversational structure to narrow the search space and lexical-semantic relevance to rank candidates. The system operates through prompting and tool use without memory-specific policy training. EGMEMORY achieves 68.2% on GroupMemBench and 77.9% on EverMemBench, outperforming the strongest evaluated baselines by 22.7 and 21.4 percentage points, respectively. It further reaches 73.6% on the dyadic LoCoMo benchmark, demonstrating generalization beyond multi-actor conversations. We will release the codebase upon formal publication.

Every Monday
Get next week’s papers.
Subscribe on Substack