Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

Egor Pakhomov and Erik Nijkamp (Salesforce AI Research) test what MemoryAgentBench's Conflict Resolution split measures by running its own stated rule, newest statement wins, as a zero-learning resolver. Accepted at the NeurIPS 2026 Interpreting Agent Behavior workshop.
Ask this paper
Baseline. The frozen last-write resolver answers 80.25% of questions under the official metric (74.5% on the three held-out fact lists), with no memory learning at all.
Unreachable gold. 67 items have a released gold answer that the last-write graph cannot reach but overwritten statements would, for example gold New Delhi after the capital is restated as Grosseto. These are a third of the multi-hop questions at 262K.
Agents on the split. Two long-context models and a re-implemented BM25 agent score 84.7%, 82.6% and 41.6% on items the rule solves, but 10.4%, 11.9% and 6.0% on the 67 unreachable ones.
Reading. Most failures are a reachability split plus a small parser-scope residual, so low scores do not by themselves show missing selective-forgetting ability.
Recommendation. Scores on this split should be read per item, not as an aggregate.
Abstract
MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.