🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Ansuman Mullick and Eray Tuzun (Bilkent University) classify personal facts into a behavioral ontology and apply category-specific retention policies as deterministic functions over LLM-extracted metadata, then locate through ablation which half of the design produces which gain.

121Memory
KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

Xi Shi and Qian Lou (University of Central Florida) build KVShareArena, a benchmark for reusing KV caches when the reused text is not a prompt prefix, which is the case for retrieval-augmented servers and for multi-agent coordinators reading reports written by other agents.

122Memory
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

Benjamin Gruenbaum, Doron Porat, Assaf Natanzon and colleagues at Eon generate a complete fictional company, including simulators of Salesforce, Zendesk, Slack and Gong, so enterprise agent answers can be graded exactly against computed answer keys.

123Evaluation
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

Pujun Zheng at East China Normal University with Shanghai Artificial Intelligence Laboratory audits SWE-Bench Pro, finds reward hacking through gold-solution leakage and task-quality defects, and releases SWE-Bench Pro Verified, on which several models score substantially lower than previously reported.

124Code
SkillAdam: Stable and Efficient Skill Evolution for Agents

SkillAdam: Stable and Efficient Skill Evolution for Agents

Gaoyuan Li, Meihao Fan, Shaolei Zhang and Ju Fan at Renmin University present SkillAdam, which ports Adam's two moment estimates to the optimization of discrete, non-differentiable skill documents so that skill self-evolution stops oscillating.

125Agents
Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Maximilian Schall, Sedigheh Eslami, Antoine Chaffin and colleagues at Perplexity AI release Q2D-Web, a 190M-document web corpus with 70k agent-reformulated search queries in ten languages, built because production RAG retrievers serve machine-written queries and existing benchmarks test human-written ones.

126Retrieval
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Sihan Ge and colleagues at Cardinal Operations and Shanghai Jiao Tong University benchmark whether an LLM agent knows when to ask a clarifying question before turning a natural-language operations research request into a mathematical program.

127Agents
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Minji Kim, Jihyoung Jang and Hyounghun Kim at POSTECH argue that non-compliance in vision-language models is evaluated at the wrong granularity, and build a benchmark where a single query mixes answerable content with content that should be withheld.

128Multimodal
Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

Peng Cui, Heejin Do and Mrinmaya Sachan at ETH Zurich apply Knowledge Space Theory, which formalizes the idea that mastering a concept requires mastering its prerequisites, as a normative standard for LLM mathematical knowledge, and compare eight models against real human learners.

129Evaluation
Single-Query Black-Box Calibration Auditing via Logit Bias

Single-Query Black-Box Calibration Auditing via Logit Bias

Roman Plaud and colleagues at Institut Polytechnique de Paris, Onepoint and Ghent University show that a logit_bias parameter is enough to recover exact probability thresholds from an API that hides output probabilities, using one query per sample.

130Evaluation
Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

Daan Henselmans, Derck Prinzhorn and Arno Libert at the Aithos Research Foundation score how well a model can defend its verdict under critical questioning, using a standard that does not require ground truth about the right answer.

131Evaluation
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Shenxi Wu and colleagues at The Chinese University of Hong Kong and Shanghai AI Laboratory build a benchmark that tests scientific document understanding as a research-assistant workflow rather than as isolated perception, retrieval and reasoning tasks, and ship the training data to improve it.

132Evaluation
Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Matthias Busch and colleagues at Helmholtz-Zentrum Hereon and Hamburg University of Technology audit 22 frontier models on 12 molecular regression benchmarks to separate models that predict a property from models that reproduce a published number.

133Retrieval
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

Abhishek Sharma builds an executable benchmark for agents resolving payment exceptions when a merchant's processor, ledger, ERP and bank feed hold contradictory beliefs about the same order, and grades on executed monetary effects rather than answer accuracy.

134Evaluation
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

Aziz Ben Amor and colleagues at Pi School release RefactorPlatform, an evaluation harness that holds the environment fixed and varies one coding-agent design axis at a time on 100 repository-scale RefactorBench tasks.

135Agents
TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

Zhibo Yang and colleagues build a benchmark for scientific discovery rather than reproduction: agents see a neutral objective and frozen data, with the source study's conclusions, expected values, and analysis path withheld.

136Evaluation
$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction

$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction

Quan Shi, Keshav Dhandhania, Karthik Narasimhan and Victor Barres at Sierra and Princeton make agent construction itself the benchmark task: a developer agent must deliver a working customer-service agent under the conditions of a real client engagement.

137Agents
Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency

Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency

Sverrir Thorgeirsson, Theo B. Weidmann and Zhendong Su at ETH Zurich run a preregistered cross-sectional study of 100 tertiary-level students to find out which measured skills predict how well someone performs at vibe coding.

138Evaluation
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Ji Soo Lee and colleagues at Meta and KAIST build WearableQA from the wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements each.

139Evaluation
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

Xinran Zhang and colleagues at GAIR ask whether enterprise-agent rankings survive a change in who the agent is competing against, and find they largely do not.

140Agents
Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

Ross Tieman and Evan Markou argue that semantic similarity is the wrong diversity measure for populations of language models, and use compression distance between raw outputs to recover the structure that predicts correlated failure.

141Efficiency
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

Axel Ahlqvist and colleagues at the UK AI Security Institute, Meridian and Anthropic attack evaluation awareness, the problem that capable models can tell when they are being tested rather than deployed, which weakens any conclusion a safety evaluation supports.

142Evaluation
Evaluating and Improving LLM Self-Modeling

Evaluating and Improving LLM Self-Modeling

Siqi Zeng, Andre N. Assis and Rowan Wang, working through the Anthropic Fellows Program, measure whether a model can answer verifiable questions about its own behavior, and then try to train the ability in.

143Evaluation
Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating

Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating

Heejin Choi at Yonsei University shows that the ranking of continual knowledge-updating methods reverses depending on when you evaluate and how much adapter capacity the baseline gets.

144Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026