🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
188 papers · RetrievalClear filters →
Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Reza Esfandiarpoor, Radek Osmulski, Even Oldridge and colleagues at NVIDIA (with the University of Edinburgh) measure what a ReAct retrieval agent adds over dense retrieval on complex search, and what it costs.

01Retrieval
EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

Zeyu Liu, Jian Zhong and colleagues led by Nankai University (EMNLP 2026 Findings) introduce EvalMem, a framework that attributes long-term memory failures to encoding, retrieval or generation instead of reporting only end-to-end QA accuracy.

02Memory
OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling

OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling

Ruiqing Zhao, Yuan Zuo and colleagues at Beihang University (EMNLP 2026 Main) introduce OptiSkill, which builds an evolving library of solver-verified formulation skills for LLMs that translate word problems into mathematical programs.

03Retrieval
When Does Execution Provenance Help Agent Memory Retrieval?

When Does Execution Provenance Help Agent Memory Retrieval?

Yiqi Wang, Taotao Cai and colleagues at the University of Southern Queensland, SUSTech, Jiangsu, Nanjing University and Norve Labs treat agent-memory retrieval as budgeted evidence completion and test when execution provenance helps retrieve all the evidence an answer needs.

04Retrieval
Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Hao Fu, Baiting Zhu, Minglei Chen, Yinjie Huang and Shuai Ding (Meta) describe EvoPilot, a human-gated method for running LLM-agent research loops on a production retrieval system, and report a 37-day campaign on the retrieval stack behind Video Deep Dive.

05Retrieval
AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

Cao, Zhou, Mei and colleagues (National University of Defense Technology) propose AutoViewMem, which organizes conversational long-term memory into automatically discovered, low-overlap semantic views at write time so that plain top-K retrieval returns focused evidence.

06Memory
M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

Yilun Liu, Shimin Tao, Daimeng Wei and colleagues at Huawei audit community agent-skill libraries, find almost no content in low-resource languages, and propose M-SQE, a post-retrieval scorer that picks usable skills from synthesized in-language candidates.

07Agents
Question's Gambit: The First Move Matters in Agentic Deep Search

Question's Gambit: The First Move Matters in Agentic Deep Search

Hamidi Rad, Clarke, Bagheri and colleagues (Toronto, Waterloo, UC Berkeley, McGill, Mila) show that the first retrieval call a deep research agent makes has a large effect on final accuracy, and propose Question's Gambit, a module that builds the opening context before the agent loop starts.

08Agents
Efficiently Linking Unstructured Data for Multi-step Reasoning

Efficiently Linking Unstructured Data for Multi-step Reasoning

Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar and Zachary Ives (University of Pennsylvania) build a query engine for the retrieval step that sits under agentic reasoning pipelines, executing filters, multi-vector search, relational joins and similarity joins together.

09Reasoning
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

Mingxuan Zhang and colleagues present RAFT, which abstracts each closed support case into a directed chain of timeline entries and retrieves at the entry level rather than treating cases as static documents.

10Retrieval
Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

Wonmi Choi and colleagues characterize the resource dynamics of LLM agents across retrieval-augmented QA, web search and software coding, and use the measurements to build two scheduling optimizations.

11Agents
AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

Z. C. Luo and a 13-author team diagnose three failures in repository-level memory retrieval for program repair, then route memory by repair stage rather than by similarity alone.

12Code
Self-Evolving Search Index

Self-Evolving Search Index

Sangam Lee and colleagues present SELF-INDEX, which lets a retrieval index diagnose its own failures, revise the responsible index keys and validate each revision before committing it.

13Retrieval
EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

Yinzhu Quan and Zefang Liu distill verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data and separate skill transfer from skill retrieval.

14Retrieval
Do Frontier Models Seek Safety Evidence Before Acting?

Do Frontier Models Seek Safety Evidence Before Acting?

Omer Tafveez (University of Michigan) introduces SAFE, a benchmark that tests whether frontier models choose to retrieve optional safety evidence before making a deployment decision, varying the evidence's retrieval cost, probability, severity and presentation.

15Safety
Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Yuanyi Song, Weinan Zhang and colleagues (SJTU, OPPO) propose REALM, an agent memory that reorganizes its graph structure based on which memories are retrieved and used together.

16Agents
On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

William L. Tong, Eran Malach, Emmanuel Abbe, Cengiz Pehlevan and colleagues (Apple, Harvard) show that the gating mechanism in state space models drives both their weak in-context retrieval and their better length generalization.

17Architecture
Prefix Sharing Is a Sorting Problem

Prefix Sharing Is a Sorting Problem

Rong He proves that choosing the order of reusable prompt pieces (retrieved passages, tool definitions, few-shot examples) to maximize prefix-cache reuse is equivalent to choosing a binary hierarchy over requests, and shows that the single global order used by deployed systems is asymptotically wrong.

18Retrieval
Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Junghyun Min (Georgetown University, as a Nokia Bell Labs intern) with Huseyin Uzunalioglu and Mohamed Trabelsi (Nokia Bell Labs) run autonomous research agents on an open-ended industrial problem, telecom ticket retrieval, and compare the outcome with 10 months of human work.

19Agents
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Utkarsh Soni and colleagues at Manulife build TAM, a benchmark of real tasks that require following manuals with tens of thousands of rules, in ICD-10-CM clinical coding and U.S. federal sentencing.

20Reasoning
MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

EunKyeong Lee, Kyeong-Jin Oh and colleagues (KT Corporation) present Mosaic, a training-free GraphRAG method in which an LLM turns each query's evidence needs into its own graph exploration policy, while the graph, indexes and answer generator stay shared.

21Retrieval
REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

Tuan Nguyen, Fan Lai and colleagues (VinUniversity and UIUC) compress RAG context by mining the generator's past attention over each document into a reusable importance store, instead of running a compressor per query.

22Retrieval
VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents

VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents

Peiyuan Gao, Wei Lu and colleagues (Renmin University of China) build VikingRAG, a directory-aware retrieval system that matches state-of-the-art structured-document RAG accuracy while using a fraction of the tokens.

23Retrieval
Kernel-Managed Shared Memory for System-Wide Personalization

Kernel-Managed Shared Memory for System-Wide Personalization

Ryan Lum and Yongfeng Zhang (Rutgers University) move memory retrieval, privacy enforcement and prompt injection out of individual agents and into the agent-system kernel, and evaluate the design on AIOS across three assistant models and 1,800 trials.

24Memory
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026