🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
188 papers · RetrievalClear filters →
KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

Xi Shi and Qian Lou (University of Central Florida) build KVShareArena, a benchmark for reusing KV caches when the reused text is not a prompt prefix, which is the case for retrieval-augmented servers and for multi-agent coordinators reading reports written by other agents.

25Memory
Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Maximilian Schall, Sedigheh Eslami, Antoine Chaffin and colleagues at Perplexity AI release Q2D-Web, a 190M-document web corpus with 70k agent-reformulated search queries in ten languages, built because production RAG retrievers serve machine-written queries and existing benchmarks test human-written ones.

26Retrieval
What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory

What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory

Chen Shen at Megagon Labs introduces the restore counterfactual, a per-question intervention that puts the gold evidence back into a reader's context after eviction, which separates losses eviction destroyed permanently from losses retrieval merely failed to surface.

27Agents
Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Matthias Busch and colleagues at Helmholtz-Zentrum Hereon and Hamburg University of Technology audit 22 frontier models on 12 molecular regression benchmarks to separate models that predict a property from models that reproduce a published number.

28Retrieval
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Ankit Goyal and Jaideep Ray at LinkedIn run a controlled study of what happens to an agent's memory store when the model reading it changes, comparing verbatim long context, chunked RAG, model-written notes, and a fixed-schema knowledge graph.

29Memory
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

Aziz Ben Amor and colleagues at Pi School release RefactorPlatform, an evaluation harness that holds the environment fixed and varies one coding-agent design axis at a time on 100 repository-scale RefactorBench tasks.

30Agents
KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

Di Chai and colleagues at Shanghai University of Finance and Economics and Peking University page an agent's overflowed workspace history as KV state across GPU memory, host memory, and NVMe, instead of compacting it into summaries or re-retrieving it as text.

31Memory
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Jiahe Geng, Jinpeng Wang and Kun Yuan build RSM-full, an online clustered-memory pipeline for LLM agents operating under a 2k to 5k prompt-token budget, and show the gain comes from how memories are merged and packed rather than from raw recall.

32Agents
STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

Retrievers chunk long documents by length, which throws away the hierarchy the document already has. Researchers at IBM point out that a table of contents already encodes that global structure, and they build a retriever around it.

33Retrieval
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Wen-Yu Chang and Yun-Nung Chen build LOCOMO-CONV, a conversational memory benchmark that replaces QA-style probing with in-situ dialog usage, and find retrieval gaps that QA benchmarks simply do not see.

34Evaluation
Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Seonghyeon Cho and Chanjun Park at Korea University show that the standard way of measuring whether agent skills help is confounded by selection bias, and introduce a matched-execution estimator that flips the conclusion for several models.

35Agents
Lazy Grounding: Attacking Search Agents with Factual Evidence

Lazy Grounding: Attacking Search Agents with Factual Evidence

Yulin Zhang and colleagues (Duke, CMU) show a search agent can be misled without any false document, by surfacing truthful evidence that answers a neighboring question instead of the one asked.

36Retrieval
Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay

Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay

Prateek Chhikara introduces matched trajectory replay, a protocol that holds answer states, evidence, budgets, and action costs fixed so confidence-to-action mappings in retrieval agents can be compared on their trajectory-level consequences rather than in isolation.

37Retrieval
Zero-Mem

Zero-Mem

Production memory stacks spend extra model calls on summarizing interactions, writing records, and reranking retrievals. Each of those calls costs tokens and latency, and the generated summaries quietly discard the evidence you later need. This work asks whether structured memory access requires generation at all.

38Memory
Progressive Disclosure, Measured

Progressive Disclosure, Measured

Agent Skills package expertise into folders an agent loads on demand, and progressive disclosure exposes only what a query needs, from a short description down to specific passages. Practitioners adopted this pattern fast for book-length tasks, but the supporting evidence was anecdotal until now.

39Agents
ReContext

ReContext

Models now support 128K context windows yet still fail to use evidence already sitting in the prompt. ReContext is a training-free inference harness for long-context reasoning that uses model-internal relevance signals to build a query-conditioned evidence pool, then replays it right before final generation while preserving the full original context.

40Memory
BlockSearch

BlockSearch

BlockSearch runs the first systematic study of in-context retrieval at the scales real retrievers actually face, million-token corpora and length generalization far beyond training size. It introduces a 0.6B language-model retriever whose architectural and training changes improve over prior LM baselines and length-generalize up to 10 times beyond their training length, pointing toward retrievers that stay reliable as context windows keep growing.

41Retrieval
Generative Skill Composition

Generative Skill Composition

Coding agents accumulate large skill libraries, and picking the right skills for a task has become the bottleneck. The usual options either dump the whole collection into context or retrieve skills with embeddings and rerankers, and both treat selection as a ranking problem rather than a joint plan. ---

42Agents
Compositional Skill Routing

Compositional Skill Routing

Real tasks rarely map to a single skill. They usually need several skills composed together, yet most skill routing still treats the problem as picking one tool from a library. This work formalizes Compositional Skill Routing, where an agent must select and sequence multiple reusable skills from large libraries to satisfy a complex query, and introduces SkillWeaver, a decompose, retrieve, and compose pipeline built around it.

43Agents
AtomMem

AtomMem

Long-term memory for LLM agents tends to fail in two ways: coarse summaries drift over time, and unconstrained updates corrupt what was already stored. AtomMem keeps the unit of memory small, using a Fact Executor that selectively extracts high-value atomic facts from long interactions and organizes them into hierarchical event structures and temporal user profiles, with an associative memory graph that reconnects fragmented memories at retrieval. The approach reports state-of-the-art results on the LoCoMo long-term memory benchmark.

44Memory
State-Externalizing Harnesses

State-Externalizing Harnesses

Harness-1 is a 20B search agent trained with reinforcement learning inside a stateful harness that offloads routine bookkeeping to the environment. The argument is that search agents are usually trained as policies over a growing transcript, forcing RL to optimize both genuine search decisions and recoverable state like which evidence is useful or which claims are checked. Harness-1 moves that state out of the policy and into an environment-side working memory of candidate pools, an importance-tagged curated set, compact evidence links, and verification records. The 20B agent reaches an average curated recall of 0.730 across eight retrieval benchmarks, beating open-source baselines by 11.4 points and matching or outperforming much larger frontier searchers, with stronger generalization on unseen domains.

45Agents
The Efficiency Frontier

The Efficiency Frontier

Context costs dominate production LLM bills, and the right strategy depends on how often preprocessing gets reused. This paper models context-strategy selection as a deployment-aware optimization problem that jointly accounts for task performance, token cost, and reuse, then uses it to compare retrieval-based and preprocessing-based approaches under realistic constraints.

46Efficiency
Memory as a Model

Memory as a Model

MeMo augments any frozen LLM with a separately trained memory model that stores, retrieves, and integrates facts on the base model's behalf. Memory updates are decoupled from base-model weight updates, so the system supports continual learning without catastrophic forgetting, a property RAG fails to deliver because a vector store is just a database with a learned encoder bolted on.

47Memory
Is Grep All You Need?

Is Grep All You Need?

The paper evaluates grep-style text search against embedding-based retrieval inside coding agents. When wrapped in a suitable agent harness, grep matches or exceeds embedding retrieval on coding-agent tasks. The study isolates the contribution of the harness from the contribution of the retrieval primitive, and finds that harness design accounts for most of the performance differential typically attributed to embeddings.

48Retrieval
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026