🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Arham Sethi and colleagues at Spark AI Research and Apta AI build a 1,024-item benchmark that forces a tool call and guarantees an unusable payload, and measure how often tool-augmented models then assert a value the tool never returned or invent a reason for withholding one.

97Agents
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

Guangsheng Yu and colleagues at the University of Technology Sydney and CSIRO build K-Bench, which scores LLM unlearning on a deployed ReAct agent by inspecting every channel where a secret can appear, and show that answer-only benchmarks overstate forgetting.

98Evaluation
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Ali Ansari, Haoran Sun and a team of physics faculty and graduate researchers led by John Sous and Arman Cohan (Yale University) audit six physics benchmarks and find that most answers graded wrong were grader errors, wrong reference solutions or ill-posed questions.

99Evaluation
GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

The standard way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. Amazon audits that gate against verifiable rewards across 25 agents from six providers, and finds two specific failures.

100Evaluation
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

The vivo AI Lab team presents BlueLM-GUI, a 35B-A3B mobile GUI agent whose data collection, RL rollouts and evaluation all run on hundreds of real phones instead of emulators.

101Agents
Scaling Clinical Judgment to Evaluate Medical AI

Scaling Clinical Judgment to Evaluate Medical AI

Thomas A. Buckley and colleagues at Harvard Medical School fine-tune PrecepTron, a 32B judge trained with LoRA on a small number of physician-scored examples, and release GRAND-ROUNDS, a benchmark of physician scores.

102Evaluation
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Utkarsh Soni and colleagues at Manulife build TAM, a benchmark of real tasks that require following manuals with tens of thousands of rules, in ICD-10-CM clinical coding and U.S. federal sentencing.

103Reasoning
Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

Carl Edwards, Gabriele Scalia and colleagues (Genentech) build AssayBench-Loop, a benchmark of 1,389 CRISPR screens for choosing experiments over multiple rounds, and AssayLoop, which combines a transformer acquisition policy trained on past screens with LLM-derived biological priors.

104Evaluation
Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

Benjamin Gruenbaum and colleagues at Eon describe a generator that builds a complete, internally consistent fictional enterprise across 66 business products with no real dataset behind it, and evaluate realism with reference-free checks fixed before tuning.

105Evaluation
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Jiaqiang Li, Tao Gui and colleagues (Fudan NLP Group with Shanghai AI Laboratory) build Sci-MMR, a benchmark that checks whether multimodal research agents recover the full evidence chain behind a scientific answer, and find answer accuracy runs more than 20 points ahead of evidence recovery.

106Multimodal
Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization

Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization

Dong-Jae Koh, Young-Kyoon Suh and colleagues (Kyungpook National University) introduce DNBENCH for database normalization from 1NF to BCNF and a multi-agent method, MARS, that improves the benchmark score by 82.0% over a single prompt.

107Agents
When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang and JinYeong Bak (Sungkyunkwan University) show that surface noise such as typos makes LLM judges report social bias that is not in the text, so bias measured on noisy text is overestimated.

108Evaluation
EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale

EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale

Ege C. Kaya and Abolfazl Hashemi (Purdue University) analyze the update rule of EGGROLL, the low-rank evolution strategy used to fine-tune LLMs without gradients, and introduce LOO-ROLL, a leave-one-out estimator that halves estimator error at equal evaluation cost.

109Evaluation
Thinking with Looped Flows

Thinking with Looped Flows

Ayhan Suleymanzade (EPFL) and colleagues at KAIST, Amsterdam, CMU, TU Wien and Oxford train looped models with local denoising objectives so that early recurrent updates learn to support later ones, and reach 58.8% on ARC-AGI-1.

110Evaluation
Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

Hanhua Hong, Chenghua Lin and colleagues (Manchester, IQuest Research, Beihang and Langboat) introduce AgentActionBench for the NLPCC 2026 shared task, which grades how agents reproduce paper experiments by recording their actions rather than only checking the final repository.

111Evaluation
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Minghao Guo, Meng Cao and colleagues (MBZUAI and USTC) build Mr.LHDR, a deep research benchmark whose questions require long chains of dependent, multimodal evidence, and find that the best system fully completes only about a third of them.

112Multimodal
The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

Makoto Fukushima, Hua-Dong Xiong and Ehsan Moradi Pari (Honda Research Institute) define the convention gap, the difference between failure rates predicted from literal messages and observed failure rates, and use Hanabi to show that human players rely on implicit conventions that AI-AI evaluation does not capture.

113Evaluation
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

Pingchen Lu, Zhongxiang Dai and colleagues (CUHK-Shenzhen, Tianjin University and NUS) treat agent skill optimization as a budgeted sequential decision problem and use a contextual bandit to decide which candidate skills are worth evaluating.

114Agents
Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

Yu-Chung Hsiao (Cisco Systems) shows that on post-2025 proprietary models, verbalized confidence is a more robust soft score for LLM-as-a-Judge than log-probabilities, which reverses the standard advice.

115Evaluation
SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

Qibai Chen (independent researcher) and Zeming Liu (Brown University) measure how well frontier LLMs resolve package version constraints across npm, PEP 440 and Cargo, and find predictable rule-specific blind spots that tool delegation removes.

116Evaluation
What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

Haseeb Mohammed Afsar (independent researcher) probes a seeded random sample of 400 servers from the 24,135-server MCP registry and finds that fewer than half start, and that popular tool-use benchmarks contain far more duplicated tools than real MCP servers do.

117Evaluation
BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

Shenghan Zheng and Christophe Hauser (Dartmouth College), with Dawn Song (UC Berkeley) and collaborators at Amazon, BenchFlow and several universities, build BenchShield, an instrumentation layer that detects reward hacking in agent benchmarks from a formal model of each evaluation's reward-relevant events.

118Evaluation
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

Zixiang Chen, Yuheng Lu and colleagues at Beihang University introduce JarvisGUI, a benchmark that evaluates GUI agents on workflows spanning Android, Windows and Ubuntu, where intermediate results must move between devices.

119Agents
Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Ansuman Mullick and Eray Tuzun (Bilkent University) classify personal facts into a behavioral ontology and apply category-specific retention policies as deterministic functions over LLM-extracted metadata, then locate through ablation which half of the design produces which gain.

120Memory
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026