🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
301 papers · MemoryClear filters →
Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies

Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies

Yi Wen, Xiangyu Zhao and colleagues at City University of Hong Kong (with Huawei) propose MemoType, which routes each memory and query to a type-specific retrieval strategy. Accepted at NeurIPS 2026.

01Memory
When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

Olukunle Owolabi, Pulkit Gupta and Fei Wang at Meta AI study when an agent's memory pipeline should store an inferred assertion, and propose confidence thresholds conditioned on the assertion's semantic category. The paper is accepted at the NeurIPS 2026 Social Agent Workshop.

02Agents
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

03Memory
Continuous Memory Machines

Continuous Memory Machines

Ciaran Regan, Kai Arulkumaran, Luke Darlow, Stefania Druga, Sebastian Risi and Llion Jones at Sakana AI introduce the Continuous Memory Machine (CMM), a recurrent architecture with separate matrix-valued short-term and long-term memory states, built on Sakana's Continuous Thought Machine. It appears at the NeurIPS 2026 PALM workshop.

04Memory
AMBER: Training Long-Horizon Web Agents through Append-Only Memory

AMBER: Training Long-Horizon Web Agents through Append-Only Memory

Chinmay Savadikar, Tianfu Wu, Lingyun Wang and colleagues at North Carolina State University and Shopify introduce AMBER, an append-only memory that a web agent learns to write while it reasons and acts, trained end-to-end with outcome-reward RL.

05Memory
MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

Hongming Xu, Zhiyu Li, Juncheng Zhang and colleagues at Shanghai Jiao Tong University and MemTensor introduce MemTrace, a memory system for long-horizon coding agents that checks recalled evidence against the current repository state before reusing it.

06Code
Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Debeshee Das (Anthropic Fellows Program) with David Huang and Javier Rando (Anthropic) show that a misaligned agent can write a goal it cannot yet act on into persistent memory, and that a later, aligned agent will often carry it out.

07Safety
Trained Agentic Context Management

Trained Agentic Context Management

Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

08Agents
Continual Graph Memory for Mathematical Research Agents

Continual Graph Memory for Mathematical Research Agents

Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and colleagues at UCLA, with senior authors including Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao and Wei Wang, present Ansatz, a mathematical research agent built around Continual Graph Memory, which stores proof progress as typed graphs instead of flat text.

09Memory
FlowState: Execution State as Memory for Long-Horizon LLM Agents

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Minghao Li, Bangyan Li and colleagues at Ant International, Ant Group propose FlowState, an agent memory framework that stores execution state as typed, linked nodes with references back to raw tool output, so an agent can revisit earlier decisions and their evidence without carrying the full history in context.

10Agents
Just-In-Time Agent Memory with Runtime Agentic Research

Just-In-Time Agent Memory with Runtime Agentic Research

Bingyu Yan, Zheng Liu and colleagues at the Beijing Academy of Artificial Intelligence propose Just-In-Time Agent Memory (JAM), which keeps complete raw histories and builds query-specific context at request time with a trained Researcher agent, instead of compressing memory before requests arrive.

11Agents
VISTA: A Visual Harness for Reasoning in an Interactive World

VISTA: A Visual Harness for Reasoning in an Interactive World

Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He at MIT introduce VISTA, a visual harness that lets a general-purpose multimodal model act in interactive environments from image observations, with a lossless visual memory it can search and rearrange while it reasons.

12Reasoning
Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Jingyuan Ma, Zhifang Sui and colleagues at ByteDance and Peking University introduce Traverse, a search harness in which the agent manages its own process through Rubric, Answer and Verify states and compresses its context with a Seal Memory tool.

13Agents
GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

Zhibang Yang and colleagues at Peking University introduce GitHarness, which stores an agent's requirement states and work states as a branchable Git-style history so that when a user adds, changes or revises a requirement, the agent restores the right earlier state instead of rewriting everything.

14Agents
ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

Zihan Wang, Xuehai Zhou and colleagues at the University of Science and Technology of China introduce ActKV, a KV cache compression method for agent inference that keeps the cache entries that matter most for generating actions.

15Memory
Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

Jiaqi Ding and Guorong Wu at UNC-Chapel Hill (EMNLP 2026 Main) introduce MemProbe, a diagnostic suite that borrows four experimental paradigms from human memory research to test how agent memory systems update, preserve and attribute information over time.

16Memory
EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

Zeyu Liu, Jian Zhong and colleagues led by Nankai University (EMNLP 2026 Findings) introduce EvalMem, a framework that attributes long-term memory failures to encoding, retrieval or generation instead of reporting only end-to-end QA accuracy.

17Memory
RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

Fanyu Zhao, Yinsheng Li and colleagues at Fudan University and the Qwen Business Unit of Alibaba introduce RPMem, a parametric memory for agents that compiles each session into a model-independent latent memory and maps it to LoRA weights for whichever backbone is in use.

18Memory
Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations

Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations

Zihao Lu, Zhihang Yuan and Lei Shi at Alibaba Cloud Computing introduce EGMemory, a memory system for long multi-party conversations that stores message-level evidence separately from an explicit, revisable state.

19Memory
ChipMEM: Verification-Grounded Memory for EDA Agents

ChipMEM: Verification-Grounded Memory for EDA Agents

Abdulrahman AlRabah and colleagues at the University of Illinois Urbana-Champaign, with a co-author from NVIDIA, build ChipMEM, a memory layer for chip-design agents that stores a skill only after an EDA tool has verified it.

20Memory
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Xinjie Shen (Georgia Tech) with Wei Fan, Dayiheng Liu and colleagues at Alibaba Token Foundry (Qwen technical report) present VHD-Play, which generates agentic RL environments by first solving a mathematical model and then rendering its decision process as stateful tools.

21Agents
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz and Shafiq Joty (Salesforce AI Research) propose Just-in-Time Memory (JitMem), which stores raw trajectories and decides what to extract from them only when a new task arrives.

22Memory
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

23Agents
When Does Execution Provenance Help Agent Memory Retrieval?

When Does Execution Provenance Help Agent Memory Retrieval?

Yiqi Wang, Taotao Cai and colleagues at the University of Southern Queensland, SUSTech, Jiangsu, Nanjing University and Norve Labs treat agent-memory retrieval as budgeted evidence completion and test when execution provenance helps retrieve all the evidence an answer needs.

24Retrieval
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026