AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Yu Lin and colleagues at AutoArk present Edge0, a streaming MoE inference engine that serves a 35B-class MoE from SSD on a single 24GB machine by predicting the next layer's expert routing one token ahead and using that prediction as the routing.

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents
Yuanyi Song, Weinan Zhang and colleagues (SJTU, OPPO) propose REALM, an agent memory that reorganizes its graph structure based on which memories are retrieved and used together.

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
Srikanta Datta Tumkur and colleagues (Vizuara) simulate KV cache placement across GPU, CPU and SSD for chat, agent and document QA sessions and find that tier capacity, not placement policy, produces the gains.

Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails
Harish Gaggar (Intuit Credit Karma) compares five context-trimming strategies for multi-step agent workflows and finds that preserving protocol-critical state matters more than the amount of text removed.

Where Should a Document Live: Context, Representations, or Parameters?
Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias and Adrià de Gispert (Amazon AGI) compare KV-cache and parametric ways of giving a model a document, across five knowledge-intensive benchmarks at matched storage budgets.

Register Tokens for Bounded-State Reasoning in Diffusion Language Models
Albert Ge, Chandan Singh, Jianfeng Gao, Frederic Sala and colleagues (Microsoft Research, UW-Madison) let masked diffusion LMs continue reasoning after earlier text is cleared by carrying state in a few register tokens.

EchoPath: Execution-Level Replayable Memory for GUI Agents
Yao Zhao and Yanxun Xu (Johns Hopkins) with Aditya Shanmugham and Swastik Roy (Amazon AGI) present EchoPath, which turns validated GUI trajectories into parameterized callable memories that replay without a fresh plan-ground-act loop.

AgentKV: Phase-Aware KV Eviction for Agentic LLMs
Taowen Tony Liu and colleagues at Imperial College London show that KV-cache eviction methods built for chat fail on agentic workloads because future queries come from distinct think, act and tool phases, and propose AgentKV, which scores cached keys against a small query buffer for each phase.

Do Not Restart: Residual Completion for Stateful Agent Handoffs
Runzhi Deng and colleagues at Nanjing University and Singapore Management University treat handing a partly finished tool-agent task from one model to another as commitment-constrained residual completion, and introduce CFRC, which lets the successor finish only the remaining work without redoing or contradicting accepted steps.

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents
Shuhuai Huang, Jingfeng Zhang and Hong Jia (University of Auckland and Fudan University) present PMPA, an attack that hides instructions in ordinary external content so that a harness-based agent writes them into its own persistent memory, where they trigger malicious actions and privacy leaks in later sessions.

MAPLE: Memory-Augmented Planning with Language and Evolution
Kesheng Chen, Yamin Hu and Wenjian Luo (Harbin Institute of Technology, Shenzhen) build MAPLE, an optimization agent that keeps an executable model of the problem and updates it across successive natural-language change requests.

FlexComp: One Model for Every Ratio in Context Compression
Kaiyan Zhao, Zhongtao Miao, Akiko Aizawa and Yoshimasa Tsuruoka (The University of Tokyo and National Institute of Informatics) train one soft context compressor that works at any compression ratio and choose the ratio for each input, instead of training a separate model for every fixed ratio.

Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
Joseph Kanichai, Tiziano De Matteis (Vrije Universiteit Amsterdam) and Animesh Trivedi (IBM Research) measure when loading KV cache from CPU or NVMe is faster than recomputing it in vLLM, and build py-kvcache, an offload connector that starts disk reads while requests are still queued.

Memory Compression for High-Fanout Agent Sandboxes
Mengming Li, Ceyu Xu and colleagues (HKUST) build AgentZip, a memory compression system for agent workloads that spawn many concurrent sandboxes from a shared template, and cut sandbox memory by up to 8.7x.

Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
Susheel Suresh and colleagues at Microsoft give the memory-curator agent in a GitHub Copilot harness read-only tools to check candidate memories against the live environment before they are saved, which roughly doubles pass rate on a database exploration benchmark and halves cost.

Kernel-Managed Shared Memory for System-Wide Personalization
Ryan Lum and Yongfeng Zhang (Rutgers University) move memory retrieval, privacy enforcement and prompt injection out of individual agents and into the agent-system kernel, and evaluate the design on AIOS across three assistant models and 1,800 trials.

ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations
Jianjie Zheng, Peng Lai, Sijie Cheng and Guanhua Chen (SUSTech, Tsinghua, RayNeo, Deepexi) propose ROAM, which classifies each incoming-versus-stored memory pair by semantic relation before deciding what to store, instead of asking an LLM to add, update, delete or rewrite in one step.

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
Xi Shi and Qian Lou (University of Central Florida) build KVShareArena, a benchmark for reusing KV caches when the reused text is not a prompt prefix, which is the case for retrieval-augmented servers and for multi-agent coordinators reading reports written by other agents.

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
Ansuman Mullick and Eray Tuzun (Bilkent University) classify personal facts into a behavioral ontology and apply category-specific retention policies as deterministic functions over LLM-extracted metadata, then locate through ablation which half of the design produces which gain.

MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging
Junxi Wang and collaborators across Shanghai Jiao Tong University, Fudan, Nanjing University, HIT and Sichuan University present MemForest, a memory compression layer that partitions history into event units, merges redundant nodes along a maximum spanning tree, and retrieves by propagating from anchor nodes.

What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory
Chen Shen at Megagon Labs introduces the restore counterfactual, a per-question intervention that puts the gold evidence back into a reader's context after eviction, which separates losses eviction destroyed permanently from losses retrieval merely failed to surface.

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course
Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta, Shashanka Ubaru and Malgorzata Zimon at IBM Research name the consistency gap, the difference between an agent's average pass rate and how often it succeeds on all five repeats of the same task, and close part of it with targeted episodic memory.

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Sequential memory agents read long documents one chunk at a time while carrying a compact memory state. That design ties reasoning depth to how far the agent has read, makes accuracy sensitive to where the evidence sits, and grows latency linearly with document length. PARSER separates reading from reasoning.

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
Chao Yao and colleagues formalize what deleting a memory record fails to do for a long-running agent, and prove how much recomputation exact forgetting requires.