VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

Liyang Fan, Min Yang, Jieping Ye and colleagues at SIAT (Chinese Academy of Sciences), SUAT and Alibaba build VibeMemBench to measure whether memory systems improve coding agents on executable repository tasks, and find that current systems mostly do not.
Ask this paper
Benchmark. 111 coding targets from 90 SWE-rebench V2 repositories plus 3,634 history trajectories from the same repositories, covering bug fixes, features, interface changes and configuration work.
Verified useful experience. A target is kept only if injecting prior experience improved its test outcome in a reference setting, so each target has experience known to help.
Direct injection helps. Transferring the verified experience to five held-out solvers raises resolution on four of them by 1.1 to 4.5 points and cuts agent steps on all five.
Memory systems fall short. When four existing memory systems must build and retrieve experience from the same history, 11 of 12 solver and system pairings fail to beat the memory-off baseline.
Takeaway. Repository history contains useful experience, and the gap lies in how memory systems extract and retrieve it.
Abstract
Coding agents operate on real repository coding tasks, and persistent memory systems promise to reuse experience across tasks. Yet existing evaluations do not show whether those systems improve executable repository work. Repository benchmarks test code changes but do not isolate memory, while memory benchmarks score recall without measuring downstream coding outcomes. We introduce VibeMemBench, a benchmark for evaluating memory systems on 111 coding targets from 90 SWE-rebench V2 repositories and 3,634 history trajectories from the target repositories. The targets follow the SWE benchmark style and cover bug fixes, feature requests, interface changes, and configuration work. An agent edits each target codebase under a declared memory condition. Executable tests decide task resolution. Each target is retained only when injected history experience improves its executable outcome in a reference setting, so every target carries a prior experience whose usefulness is verified by execution in that setting. The frozen verified experience is then transferred to five held-out solvers. Direct injection raises observed task resolution on four of them by 1.1 to 4.5 percentage points while lowering agent steps on all five. Yet when four existing memory systems must construct and retrieve experience from the same history, eleven of twelve solver and system pairings fail to exceed the matched memory-off baseline. VibeMemBench exposes the gap between the useful experience that repository history holds and the experience existing memory systems deliver for repository coding tasks.