AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
Reza Esfandiarpoor, Radek Osmulski, Even Oldridge and colleagues at NVIDIA (with the University of Edinburgh) measure what a ReAct retrieval agent adds over dense retrieval on complex search, and what it costs.

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems
Zeyu Liu, Jian Zhong and colleagues led by Nankai University (EMNLP 2026 Findings) introduce EvalMem, a framework that attributes long-term memory failures to encoding, retrieval or generation instead of reporting only end-to-end QA accuracy.

OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling
Ruiqing Zhao, Yuan Zuo and colleagues at Beihang University (EMNLP 2026 Main) introduce OptiSkill, which builds an evolving library of solver-verified formulation skills for LLMs that translate word problems into mathematical programs.

When Does Execution Provenance Help Agent Memory Retrieval?
Yiqi Wang, Taotao Cai and colleagues at the University of Southern Queensland, SUSTech, Jiangsu, Nanjing University and Norve Labs treat agent-memory retrieval as budgeted evidence completion and test when execution provenance helps retrieve all the evidence an answer needs.

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale
Hao Fu, Baiting Zhu, Minglei Chen, Yinjie Huang and Shuai Ding (Meta) describe EvoPilot, a human-gated method for running LLM-agent research loops on a production retrieval system, and report a 37-day campaign on the retrieval stack behind Video Deep Dive.

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
Cao, Zhou, Mei and colleagues (National University of Defense Technology) propose AutoViewMem, which organizes conversational long-term memory into automatically discovered, low-overlap semantic views at write time so that plain top-K retrieval returns focused evidence.

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use
Yilun Liu, Shimin Tao, Daimeng Wei and colleagues at Huawei audit community agent-skill libraries, find almost no content in low-resource languages, and propose M-SQE, a post-retrieval scorer that picks usable skills from synthesized in-language candidates.

Question's Gambit: The First Move Matters in Agentic Deep Search
Hamidi Rad, Clarke, Bagheri and colleagues (Toronto, Waterloo, UC Berkeley, McGill, Mila) show that the first retrieval call a deep research agent makes has a large effect on final accuracy, and propose Question's Gambit, a module that builds the opening context before the agent loop starts.

Efficiently Linking Unstructured Data for Multi-step Reasoning
Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar and Zachary Ives (University of Pennsylvania) build a query engine for the retrieval step that sits under agentic reasoning pipelines, executing filters, multi-vector search, relational joins and similarity joins together.

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Mingxuan Zhang and colleagues present RAFT, which abstracts each closed support case into a directed chain of timeline entries and retrieves at the entry level rather than treating cases as static documents.

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
Wonmi Choi and colleagues characterize the resource dynamics of LLM agents across retrieval-augmented QA, web search and software coding, and use the measurements to build two scheduling optimizations.

AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair
Z. C. Luo and a 13-author team diagnose three failures in repository-level memory retrieval for program repair, then route memory by repair stage rather than by similarity alone.

Self-Evolving Search Index
Sangam Lee and colleagues present SELF-INDEX, which lets a retrieval index diagnose its own failures, revise the responsible index keys and validate each revision before committing it.

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
Yinzhu Quan and Zefang Liu distill verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data and separate skill transfer from skill retrieval.

Do Frontier Models Seek Safety Evidence Before Acting?
Omer Tafveez (University of Michigan) introduces SAFE, a benchmark that tests whether frontier models choose to retrieve optional safety evidence before making a deployment decision, varying the evidence's retrieval cost, probability, severity and presentation.

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents
Yuanyi Song, Weinan Zhang and colleagues (SJTU, OPPO) propose REALM, an agent memory that reorganizes its graph structure based on which memories are retrieved and used together.

On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models
William L. Tong, Eran Malach, Emmanuel Abbe, Cengiz Pehlevan and colleagues (Apple, Harvard) show that the gating mechanism in state space models drives both their weak in-context retrieval and their better length generalization.

Prefix Sharing Is a Sorting Problem
Rong He proves that choosing the order of reusable prompt pieces (retrieved passages, tool definitions, few-shot examples) to maximize prefix-cache reuse is equivalent to choosing a binary hierarchy over requests, and shows that the single global order used by deployed systems is asymptotically wrong.

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
Junghyun Min (Georgetown University, as a Nokia Bell Labs intern) with Huseyin Uzunalioglu and Mohamed Trabelsi (Nokia Bell Labs) run autonomous research agents on an open-ended industrial problem, telecom ticket retrieval, and compare the outcome with 10 months of human work.

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni and colleagues at Manulife build TAM, a benchmark of real tasks that require following manuals with tens of thousands of rules, in ICD-10-CM clinical coding and U.S. federal sentencing.

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG
EunKyeong Lee, Kyeong-Jin Oh and colleagues (KT Corporation) present Mosaic, a training-free GraphRAG method in which an LLM turns each query's evidence needs into its own graph exploration policy, while the graph, indexes and answer generator stay shared.

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving
Tuan Nguyen, Fan Lai and colleagues (VinUniversity and UIUC) compress RAG context by mining the generator's past attention over each document into a reusable importance store, instead of running a compressor per query.

VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents
Peiyuan Gao, Wei Lu and colleagues (Renmin University of China) build VikingRAG, a directory-aware retrieval system that matches state-of-the-art structured-document RAG accuracy while using a fraction of the tokens.

Kernel-Managed Shared Memory for System-Wide Personalization
Ryan Lum and Yongfeng Zhang (Rutgers University) move memory retrieval, privacy enforcement and prompt injection out of individual agents and into the agent-system kernel, and evaluate the design on AIOS across three assistant models and 1,800 trials.