AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
MedQA-MM: Shortcuts Behind Medical Visual Reasoning
Benlu Wang and colleagues at UMass Amherst and Yale separate the answer from the route that produced it in medical multimodal MCQs, and find that scores substantially overstate image reasoning.

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
Yan Tang and colleagues formalize proactive service as a partially observable sequential decision process constrained by authorization and risk, where staying silent is a first-class action with option value.

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
Yaxing Lyu and colleagues build KC-Bench to measure whether a tool-using model can reconcile user instructions, its own parametric knowledge and live environmental observations before it acts on any of them.

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Haoyuan Zhu at the University of Sheffield with Ranplan Wireless and Cambridge AI+ preregisters a reliability study of black-box LLM observers on shared serving endpoints and reports that the instrument itself is unstable enough to invalidate gates built on it.

FailBench: How Reliable are VLMs at Judging Robot Task Success?
Zaruhi Navasardyan, Tatul Danielyan and Hrant Davtyan at Metric AI Lab assemble 2,197 real manipulation attempts from 14 sources and find that vision-language models used as robot success detectors reach only 0.77 mean balanced accuracy, with fine-tuned detectors doing worse than general-purpose models.

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
Wen-Yu Chang and Yun-Nung Chen build LOCOMO-CONV, a conversational memory benchmark that replaces QA-style probing with in-situ dialog usage, and find retrieval gaps that QA benchmarks simply do not see.

PatchBench: Evaluating AI Agents for Vulnerability Patching
Chihao Shen and colleagues at Maryland and UC Davis show that PoC-only validation inflates vulnerability-patching solve rates by 1.83x on average, because agents either recall the historical developer patch or fix the crash rather than the bug.

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
Xin He and colleagues at Sun Yat-sen University introduce SWE-Gate, a repository-level benchmark that scores coding agents on review-derived acceptance constraints alongside functional tests, and shows that passing the tests is far from passing review.

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Austin Tudor David Andrews, Jakob Foerster, Rui Ponte Costa and colleagues (Oxford, Google DeepMind, UK AI Security Institute) release CivBench, an open-source benchmark that drives language agents through 300+ turn games of Civilization VI over 76 MCP tools, and report two behavioral failures that are more interesting than the scores.

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
Vansh Wahi reports months of running autonomous prompt-optimization loops in production across contract analysis, compliance review and code quality, and catalogs eleven distinct ways the evaluation signal failed.

READY or Not: Reliable Enterprise Agent Deployment
Veronica Chatrath, Yuan Xue and a Scale AI team introduce READY, a framework that stops asking how well an agent performs and starts asking under what oversight policy and at what cost it can be deployed at a required reliability level.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
Xincheng Wei and colleagues at Meituan show that the direction a self-play curriculum needs can be derived from the solver's own failure history rather than from external task resources or generic difficulty signals.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
Yuhao Wu and a large multi-institution team introduce HarnessDev, which moves the unit of evaluation from a model's task outputs to the runnable agent harness it can build and then improve, across creation and evolution stages.

Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation
Will Badr asks whether a hint that turns a failing program into a passing one supplies missing information or merely steers the model to a solution it could already reach, and finds mostly the latter.

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
Fanrui Zhang and a large Alibaba-affiliated team propose ARISE-RL, a co-evolutionary loop in which a task and rubric Generator and a reasoning Solver train each other, replacing the verifiable gold answer that open-ended agentic RL does not have.

Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR
Esther Xin audits the verifier rather than the model, applying metamorphic testing across 307,420 verdicts from four widely used RLVR verifiers to decompose exactly which answer forms consume the error budget.

VoiceLongMemEval: Do Assistants Remember How You Sounded?
Ramit Pahwa, Parivesh Priye, and Apoorva Beedu build VoiceLongMemEval, a long-horizon conversational memory benchmark where every answer depends on how something was said rather than on what was said.

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents
Seonghyeon Cho and Chanjun Park at Korea University show that the standard way of measuring whether agent skills help is confounded by selection bias, and introduce a matched-execution estimator that flips the conclusion for several models.

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Agent benchmarks usually end when the session does. The Qwen team built one that runs an agent through a simulated 365-day year operating several online stores at once, then scored 18 frontier models on seven dimensions.

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
Pradyumna Shyama Prasad and colleagues plant optional shortcuts inside ML tasks themselves and find that 57.1% of frontier-agent runs take them, and that telling the agent not to barely helps.

AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Zixiang Xu, Jiaan Wang and Fandong Meng (WeChat AI) turn combinatorial optimization problems into partially observed tool-use environments with certified global optima, and find leading models reach exact optimality only 38.61% of the time.

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
Gyuhyeong Kim and colleagues characterize the gap between curated GitHub issues and real user requests, then build 381 multi-variant task families from SWE-bench Verified and Pro that hold the gold patch fixed while varying information composition and linguistic style.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
Yi Wang and colleagues (AMAP / Alibaba with collaborators) benchmark the outer loop rather than the coding agent, evaluating a Controller model that instructs a fixed Worker coding agent after each round and decides when to stop.

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence
Ante Kapetanovic and colleagues run 192,000 evaluations to show that putting a prior score in a judge's context metadata drags its rating toward that number, breaking the independence assumption every refinement pipeline relies on.