🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
MedQA-MM: Shortcuts Behind Medical Visual Reasoning

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Benlu Wang and colleagues at UMass Amherst and Yale separate the answer from the route that produced it in medical multimodal MCQs, and find that scores substantially overstate image reasoning.

145Evaluation
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

Yan Tang and colleagues formalize proactive service as a partially observable sequential decision process constrained by authorization and risk, where staying silent is a first-class action with option value.

146Agents
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

Yaxing Lyu and colleagues build KC-Bench to measure whether a tool-using model can reconcile user instructions, its own parametric knowledge and live environmental observations before it acts on any of them.

147Evaluation
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Haoyuan Zhu at the University of Sheffield with Ranplan Wireless and Cambridge AI+ preregisters a reliability study of black-box LLM observers on shared serving endpoints and reports that the instrument itself is unstable enough to invalidate gates built on it.

148Evaluation
FailBench: How Reliable are VLMs at Judging Robot Task Success?

FailBench: How Reliable are VLMs at Judging Robot Task Success?

Zaruhi Navasardyan, Tatul Danielyan and Hrant Davtyan at Metric AI Lab assemble 2,197 real manipulation attempts from 14 sources and find that vision-language models used as robot success detectors reach only 0.77 mean balanced accuracy, with fine-tuned detectors doing worse than general-purpose models.

149Evaluation
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Wen-Yu Chang and Yun-Nung Chen build LOCOMO-CONV, a conversational memory benchmark that replaces QA-style probing with in-situ dialog usage, and find retrieval gaps that QA benchmarks simply do not see.

150Evaluation
PatchBench: Evaluating AI Agents for Vulnerability Patching

PatchBench: Evaluating AI Agents for Vulnerability Patching

Chihao Shen and colleagues at Maryland and UC Davis show that PoC-only validation inflates vulnerability-patching solve rates by 1.83x on average, because agents either recall the historical developer patch or fix the crash rather than the bug.

151Agents
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Xin He and colleagues at Sun Yat-sen University introduce SWE-Gate, a repository-level benchmark that scores coding agents on review-derived acceptance constraints alongside functional tests, and shows that passing the tests is far from passing review.

152Code
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

Austin Tudor David Andrews, Jakob Foerster, Rui Ponte Costa and colleagues (Oxford, Google DeepMind, UK AI Security Institute) release CivBench, an open-source benchmark that drives language agents through 300+ turn games of Civilization VI over 76 MCP tools, and report two behavioral failures that are more interesting than the scores.

153Agents
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

Vansh Wahi reports months of running autonomous prompt-optimization loops in production across contract analysis, compliance review and code quality, and catalogs eleven distinct ways the evaluation signal failed.

154Evaluation
READY or Not: Reliable Enterprise Agent Deployment

READY or Not: Reliable Enterprise Agent Deployment

Veronica Chatrath, Yuan Xue and a Scale AI team introduce READY, a framework that stops asking how well an agent performs and starts asking under what oversight policy and at what cost it can be deployed at a required reliability level.

155Agents
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

Xincheng Wei and colleagues at Meituan show that the direction a self-play curriculum needs can be derived from the solver's own failure history rather than from external task resources or generic difficulty signals.

156Memory
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Yuhao Wu and a large multi-institution team introduce HarnessDev, which moves the unit of evaluation from a model's task outputs to the runnable agent harness it can build and then improve, across creation and evolution stages.

157Agents
Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

Will Badr asks whether a hint that turns a failing program into a passing one supplies missing information or merely steers the model to a solution it could already reach, and finds mostly the latter.

158Code
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

Fanrui Zhang and a large Alibaba-affiliated team propose ARISE-RL, a co-evolutionary loop in which a task and rubric Generator and a reasoning Solver train each other, replacing the verifiable gold answer that open-ended agentic RL does not have.

159Agents
Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR

Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR

Esther Xin audits the verifier rather than the model, applying metamorphic testing across 307,420 verdicts from four widely used RLVR verifiers to decompose exactly which answer forms consume the error budget.

160Evaluation
VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval: Do Assistants Remember How You Sounded?

Ramit Pahwa, Parivesh Priye, and Apoorva Beedu build VoiceLongMemEval, a long-horizon conversational memory benchmark where every answer depends on how something was said rather than on what was said.

161Evaluation
Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Seonghyeon Cho and Chanjun Park at Korea University show that the standard way of measuring whether agent skills help is confounded by selection bias, and introduce a matched-execution estimator that flips the conclusion for several models.

162Agents
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

Agent benchmarks usually end when the session does. The Qwen team built one that runs an agent through a simulated 365-day year operating several online stores at once, then scored 18 frontier models on seven dimensions.

163Agents
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

Pradyumna Shyama Prasad and colleagues plant optional shortcuts inside ML tasks themselves and find that 57.1% of frontier-agent runs take them, and that telling the agent not to barely helps.

164Agents
AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds

AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds

Zixiang Xu, Jiaan Wang and Fandong Meng (WeChat AI) turn combinatorial optimization problems into partially observed tool-use environments with certified global optima, and find leading models reach exact optimality only 38.61% of the time.

165Agents
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Gyuhyeong Kim and colleagues characterize the gap between curated GitHub issues and real user requests, then build 381 multi-variant task families from SWE-bench Verified and Pro that hold the gold patch fixed while varying information composition and linguistic style.

166Code
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Yi Wang and colleagues (AMAP / Alibaba with collaborators) benchmark the outer loop rather than the coding agent, evaluating a Controller model that instructs a fixed Worker coding agent after each round and decides when to stop.

167Evaluation
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Ante Kapetanovic and colleagues run 192,000 evaluations to show that putting a prior score in a judge's context metadata drags its rating toward that number, breaking the independence assumption every refinement pipeline relies on.

168Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026