AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
MoM: Memory of Memory
Bowen Qin and Yao Lu at the National University of Singapore propose Memory of Memory (MoM): agent memory that commits the current value on arrival while keeping the displaced value and its provenance.

Clarification Is Not Correction: LLMs Fail to Let Go
Jianzhe Lin and colleagues at Meta AI argue that many multi-turn failures come from early commitment rather than forgetting: an ambiguous first turn becomes a fixed task state that later clarification only patches.

DolphinBench: Mapping the Pareto Frontier of Agent Memory
Soumil Rathi, Deshraj Yadav and Taranjeet Singh (Mem0) release DolphinBench, a memory benchmark that scores agents on actions they take in simulated apps rather than on answers to recall questions, and requires every submission to report cost and latency with accuracy.

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
Pirzada Suhail, Menglin Xia and colleagues at Microsoft Research and M365 propose Pseudo Self-Distillation (PSD), which trains small Qwen3 models to run a multi-stage memory-construction pipeline that normally needs GPT-4.1-mini, using only the oracle's text outputs.

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks
Liyang Fan, Min Yang, Jieping Ye and colleagues at SIAT (Chinese Academy of Sciences), SUAT and Alibaba build VibeMemBench to measure whether memory systems improve coding agents on executable repository tasks, and find that current systems mostly do not.

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
Ruike Cao, Fanyu Zhao and colleagues at Alibaba's Qwen Applications Business Group, USTC and Fudan introduce MemCalib, a benchmark for whether a model gives each retrieved memory the right amount of influence on its answer, and MemCalib-RL to train that behavior.

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
Cao, Zhou, Mei and colleagues (National University of Defense Technology) propose AutoViewMem, which organizes conversational long-term memory into automatically discovered, low-overlap semantic views at write time so that plain top-K retrieval returns focused evidence.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Du, Yan, Flores and Kadav (Adobe and Brown University) keep a frontier model frozen while it operates professional design software through more than 230 tools, and let an external procedural memory of natural-language skills grow and improve from real user traffic.

CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop
He, Wu, Zhang, Zhao, He and Li (King's College London) present CoLearn, an agentic tutor that keeps an evidence-grounded memory of each learner's mastery and misconceptions and uses it to choose the next question.

LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents
Siddharth Sharma and colleagues at UC San Diego and West Virginia University introduce LIMBO, an online method that decides for each incoming task how much past experience a lifelong agent should replay into its prompt and how much inference budget to spend.

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory
Zhao and colleagues introduce a memory architecture that labels each write with its intended lifetime, so information meant to apply only to the current context cannot overwrite knowledge meant to persist.

Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits
Mingyang Mao, Wyatt Mackey and Xiaomin Lin study how to repair a reused KV cache after a document edit with a limited recomputation budget, comparing training-free choices of which positions to recompute.

CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems
Yu and colleagues propose a two-tier memory for multi-agent systems that keeps each agent's private experience separate from the group's shared knowledge, so shared memory does not erase what makes individual agents different.

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
Rao and Jaggi build a measurement harness that makes the per-call input-token budget the independent variable when comparing agent memory strategies, and report budget-violation rates as a first-class outcome rather than a footnote.

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Mingxuan Zhang and colleagues present RAFT, which abstracts each closed support case into a directed chain of timeline entries and retrieves at the entry level rather than treating cases as static documents.

LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents
Meysam Ghaffari and colleagues turn a small labeled batch of clinical coding errors into a structured mistake database and route each lesson to the agent role that can act on it.

AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair
Z. C. Luo and a 13-author team diagnose three failures in repository-level memory retrieval for program repair, then route memory by repair stage rather than by similarity alone.

JustMem: Just-Enough Memory Access for Long-Term Conversations
Guanhua Chen and colleagues formulate conversational memory access along two dimensions, discovery breadth and reading fidelity, and build JustMem to choose the right combination per query.

On-Demand Attention: Language Models Know When to Recall
Haibo Feng and colleagues show that a pretrained model's decoding states already predict whether a global attention read will help, and use that signal to invoke global attention selectively.

Reputation as Community Memory for the Agentic Web
Ryan Chard and colleagues present Cairn, a community reputation platform that lets agents query collective opinion about a data source, service or tool before using it and submit evidence-backed ratings afterwards.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-AI releases DeepSeek-V4.1-Flash, a 552B-parameter multimodal MoE built around a Causal Encoder-Decoder architecture that activates 16B parameters per decode token and 8B per prefill token, and cuts the resident KV cache to 890 bytes per token.

An Empirical Study of Harness Design for Coding Agents
Run-Ze Fan and colleagues at UMass Amherst, Emory, UNC Charlotte and Zoom hold a coding harness's execution loop fixed and vary three components (planning, action space, context management) across 176 matched settings, four models, SWE-Bench Verified and Terminal-Bench 2.1.

Agora: Git as Shared Memory for Collective AutoResearch
Yifan Zhang, Yi Dong and colleagues at NVIDIA present Agora, a shared memory for autonomous research agents in which every result, hypothesis and verification is an immutable Git commit in an append-only DAG, and report a 12-day run with 13 LLM workers and no central planner.