AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies
Yi Wen, Xiangyu Zhao and colleagues at City University of Hong Kong (with Huawei) propose MemoType, which routes each memory and query to a type-specific retrieval strategy. Accepted at NeurIPS 2026.

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory
Olukunle Owolabi, Pulkit Gupta and Fei Wang at Meta AI study when an agent's memory pipeline should store an inferred assertion, and propose confidence thresholds conditioned on the assertion's semantic category. The paper is accepted at the NeurIPS 2026 Social Agent Workshop.

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

Continuous Memory Machines
Ciaran Regan, Kai Arulkumaran, Luke Darlow, Stefania Druga, Sebastian Risi and Llion Jones at Sakana AI introduce the Continuous Memory Machine (CMM), a recurrent architecture with separate matrix-valued short-term and long-term memory states, built on Sakana's Continuous Thought Machine. It appears at the NeurIPS 2026 PALM workshop.

AMBER: Training Long-Horizon Web Agents through Append-Only Memory
Chinmay Savadikar, Tianfu Wu, Lingyun Wang and colleagues at North Carolina State University and Shopify introduce AMBER, an append-only memory that a web agent learns to write while it reasons and acts, trained end-to-end with outcome-reward RL.

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents
Hongming Xu, Zhiyu Li, Juncheng Zhang and colleagues at Shanghai Jiao Tong University and MemTensor introduce MemTrace, a memory system for long-horizon coding agents that checks recalled evidence against the current repository state before reusing it.

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough
Debeshee Das (Anthropic Fellows Program) with David Huang and Javier Rando (Anthropic) show that a misaligned agent can write a goal it cannot yet act on into persistent memory, and that a later, aligned agent will often carry it out.

Trained Agentic Context Management
Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

Continual Graph Memory for Mathematical Research Agents
Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and colleagues at UCLA, with senior authors including Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao and Wei Wang, present Ansatz, a mathematical research agent built around Continual Graph Memory, which stores proof progress as typed graphs instead of flat text.

FlowState: Execution State as Memory for Long-Horizon LLM Agents
Minghao Li, Bangyan Li and colleagues at Ant International, Ant Group propose FlowState, an agent memory framework that stores execution state as typed, linked nodes with references back to raw tool output, so an agent can revisit earlier decisions and their evidence without carrying the full history in context.

Just-In-Time Agent Memory with Runtime Agentic Research
Bingyu Yan, Zheng Liu and colleagues at the Beijing Academy of Artificial Intelligence propose Just-In-Time Agent Memory (JAM), which keeps complete raw histories and builds query-specific context at request time with a trained Researcher agent, instead of compressing memory before requests arrive.

VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He at MIT introduce VISTA, a visual harness that lets a general-purpose multimodal model act in interactive environments from image observations, with a lossless visual memory it can search and rearrange while it reasons.

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search
Jingyuan Ma, Zhifang Sui and colleagues at ByteDance and Peking University introduce Traverse, a search harness in which the agent manages its own process through Rubric, Answer and Verify states and compresses its context with a Seal Memory tool.

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements
Zhibang Yang and colleagues at Peking University introduce GitHarness, which stores an agent's requirement states and work states as a branchable Git-style history so that when a user adds, changes or revises a requirement, the agent restores the right earlier state instead of rewriting everything.

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
Zihan Wang, Xuehai Zhou and colleagues at the University of Science and Technology of China introduce ActKV, a KV cache compression method for agent inference that keeps the cache entries that matter most for generating actions.

Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
Jiaqi Ding and Guorong Wu at UNC-Chapel Hill (EMNLP 2026 Main) introduce MemProbe, a diagnostic suite that borrows four experimental paradigms from human memory research to test how agent memory systems update, preserve and attribute information over time.

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems
Zeyu Liu, Jian Zhong and colleagues led by Nankai University (EMNLP 2026 Findings) introduce EvalMem, a framework that attributes long-term memory failures to encoding, retrieval or generation instead of reporting only end-to-end QA accuracy.

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents
Fanyu Zhao, Yinsheng Li and colleagues at Fudan University and the Qwen Business Unit of Alibaba introduce RPMem, a parametric memory for agents that compiles each session into a model-independent latent memory and maps it to LoRA weights for whichever backbone is in use.

Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations
Zihao Lu, Zhihang Yuan and Lei Shi at Alibaba Cloud Computing introduce EGMemory, a memory system for long multi-party conversations that stores message-level evidence separately from an explicit, revisable state.

ChipMEM: Verification-Grounded Memory for EDA Agents
Abdulrahman AlRabah and colleagues at the University of Illinois Urbana-Champaign, with a co-author from NVIDIA, build ChipMEM, a memory layer for chip-design agents that stores a skill only after an EDA tool has verified it.

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen (Georgia Tech) with Wei Fan, Dayiheng Liu and colleagues at Alibaba Token Foundry (Qwen technical report) present VHD-Play, which generates agentic RL environments by first solving a mathematical model and then rendering its decision process as stateful tools.

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz and Shafiq Joty (Salesforce AI Research) propose Just-in-Time Memory (JitMem), which stores raw trajectories and decides what to extract from them only when a new task arrives.

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

When Does Execution Provenance Help Agent Memory Retrieval?
Yiqi Wang, Taotao Cai and colleagues at the University of Southern Queensland, SUSTech, Jiangsu, Nanjing University and Norve Labs treat agent-memory retrieval as budgeted evidence completion and test when execution provenance helps retrieve all the evidence an answer needs.