AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
Ankit Goyal and Jaideep Ray at LinkedIn run a controlled study of what happens to an agent's memory store when the model reading it changes, comparing verbatim long context, chunked RAG, model-written notes, and a fixed-schema knowledge graph.

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU
Di Chai and colleagues at Shanghai University of Finance and Economics and Peking University page an agent's overflowed workspace history as KV state across GPU memory, host memory, and NVMe, instead of compacting it into summaries or re-retrieving it as text.

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
Chao Yao and colleagues formalize what deleting a memory record fails to do for a long-running agent, and prove how much recomputation exact forgetting requires.

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory
Kazuki Nakayashiki runs twelve registered studies and 14,760 attempts on one question: when an agent inherits terse memories and can pull only one archived source record, what form of directive written into the store actually steers that choice.

SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
Qi Liu, Qinzheng Wang and Yiming Bie build SimSkill, a self-evolving agent over the SUMO traffic simulator that finds its own capability gaps, writes and solves grounded tasks, and consolidates the results into episodic, procedural and semantic memory without touching the backbone weights.

Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation
Xuanfa Jin and colleagues at CASIA and UCL attack the shared-misconception failure in multi-agent debate with R2-MAD, giving debating agents an experience memory from past debates plus per-agent confidence weights.

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory
Yuxiang Wang and colleagues fix two limitations of looped-layer latent recurrence at once, letting each iteration attend to its own earlier states and letting the model decide how many loops a given input deserves.

VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch
WenJie Fan finds that a NoPE MLA model's cache already carries a query-independent salience signal in the 64-dimensional decoupled branch, a vestige of RoPE that NoPE training repurposes, and evicts on it.

What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
Bo Zeng and colleagues show that the temporal rule aggregating KV scores across decode steps, usually treated as an implementation detail, dominates the scoring function everyone is publishing about.

SGD-KV: Summarization Guided KV Cache Compression
Zeyu Liu and colleagues present SGD-KV, which identifies attention heads specialized in hierarchical aggregation using a chunk-summarization diagnostic and allocates KV cache budget to them instead of applying a uniform heuristic.

Free Pause Tokens
John Langford and colleagues (Microsoft Research, Cornell, CMU) give a language model extra compute per next-token prediction by running it in a parallel prediction stream over a weight-shared backbone instead of spending a sequence position on it.

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Heng Wang and colleagues at Salesforce AI Research, UIUC and Cornell show that the token-importance scores every KV cache evictor computes are close to worthless: evicting uniformly at random inside each head matches the strongest prior evictor while serving 32 to 43% higher throughput.

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
Evan Chen, Shiqiang Wang and Christopher Brinton (Purdue and Exeter) name stale-plan execution, where a distributed agent team reads perfectly fresh shared state and still acts on a plan derived from a requirement that has since been superseded.

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
Wen-Yu Chang and Yun-Nung Chen build LOCOMO-CONV, a conversational memory benchmark that replaces QA-style probing with in-situ dialog usage, and find retrieval gaps that QA benchmarks simply do not see.

RuleMem: Active Rule Memory for Long-Term Conversational Agents
Xingyuan Zeng and colleagues propose RuleMem, which induces reusable natural-language Horn clauses from conversation history so that agent memory actively guides retrieval and reasoning instead of sitting as passively stored facts.

Language Models Can Control Their Own Attention
A model reads its entire KV cache on every generated token even though it ends up attending to a tiny slice of it. Ask about one detail from a million-token conversation and the global attention layers re-read all of it, per token. Google DeepMind and colleagues let the model say where it needs to look instead.

CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning
Yongshi Ye, Tian Lan and colleagues (Xiamen University, Alibaba International) propose CHIME, a self-evolving memory framework that fixes the credit assignment problem in experience memory by attributing an outcome before writing it anywhere.

Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers
Where you put a reasoning trace changes long-context accuracy by up to 50 points. Transformers process causally, so a task state discovered late cannot guide reading that already happened, and Trace as State puts the collected trace before the long-context block on a fresh pass instead of appending it after. On GraphWalks Parents, DeepSeek V4 Pro Preview goes from 29.2% on the initial pass and 43.0% with the matched append control to 81.8%, and GLM-5.2 goes from 66.4% and 83.2% to 100.0%. It wins in 26 of 27 reported combinations of model, task, and metric with no architecture change.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
Xincheng Wei and colleagues at Meituan show that the direction a self-play curriculum needs can be derived from the solver's own failure history rather than from external task resources or generic difficulty signals.

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
We describe an agent by whatever model and harness it happens to run on, which works for one session and says very little about an agent running for months across a new model, a new harness, or a new machine. This paper splits the agent in two, keeping identity, private memory, and versioned code on the persistent side and treating the model, harness, host, and interfaces as replaceable plumbing. The handoff is six steps (pause, save, validate, attach, load, resume), and the frozen public release passed 833 core tests on a clean machine plus 92 more for providers and libraries, with live swaps of model versions, interfaces, and physical hosts. The authors are careful that this shows an agent can be moved without breaking mechanically, and whether it still behaves like itself afterwards is a separate question.

MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence
Walid Saidi closes the publication gap left by MutMem V1 with a full portable verification contract for cryptographically authorized mutation of persistent agent memory, specifying canonical bytes, commitments, revocation, and a clean-install reproduction path.

Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
Jinqing Zhao and Chengcan Wu argue that prospective memory, carrying out a deferred intention at the right future cue, is schema-constrained state tracking rather than open-ended reasoning, and show that typing the action space lets small models beat the published large-model scaffold.

VoiceLongMemEval: Do Assistants Remember How You Sounded?
Ramit Pahwa, Parivesh Priye, and Apoorva Beedu build VoiceLongMemEval, a long-horizon conversational memory benchmark where every answer depends on how something was said rather than on what was said.

Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers
Egor Pakhomov and Erik Nijkamp treat a long-horizon agent's trace as a shared resource with two consumers, the human watching the run and the agent whose bounded context the trace must fold back into, and build an append-only event ledger compiled into per-consumer views.