AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
Arham Sethi and colleagues at Spark AI Research and Apta AI build a 1,024-item benchmark that forces a tool call and guarantees an unusable payload, and measure how often tool-augmented models then assert a value the tool never returned or invent a reason for withholding one.

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
Guangsheng Yu and colleagues at the University of Technology Sydney and CSIRO build K-Bench, which scores LLM unlearning on a deployed ReAct agent by inspecting every channel where a secret can appear, and show that answer-only benchmarks overstate forgetting.

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Ali Ansari, Haoran Sun and a team of physics faculty and graduate researchers led by John Sous and Arman Cohan (Yale University) audit six physics benchmarks and find that most answers graded wrong were grader errors, wrong reference solutions or ill-posed questions.

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
The standard way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. Amazon audits that gate against verifiable rewards across 25 agents from six providers, and finds two specific failures.

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
The vivo AI Lab team presents BlueLM-GUI, a 35B-A3B mobile GUI agent whose data collection, RL rollouts and evaluation all run on hundreds of real phones instead of emulators.

Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley and colleagues at Harvard Medical School fine-tune PrecepTron, a 32B judge trained with LoRA on a small number of physician-scored examples, and release GRAND-ROUNDS, a benchmark of physician scores.

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni and colleagues at Manulife build TAM, a benchmark of real tasks that require following manuals with tens of thousands of rules, in ICD-10-CM clinical coding and U.S. federal sentencing.

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens
Carl Edwards, Gabriele Scalia and colleagues (Genentech) build AssayBench-Loop, a benchmark of 1,389 CRISPR screens for choosing experiments over multiple rounds, and AssayLoop, which combines a transformer acquisition policy trained on past screens with LLM-derived biological priors.

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data
Benjamin Gruenbaum and colleagues at Eon describe a generator that builds a complete, internally consistent fictional enterprise across 66 business products with no real dataset behind it, and evaluate realism with reference-free checks fixed before tuning.

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Jiaqiang Li, Tao Gui and colleagues (Fudan NLP Group with Shanghai AI Laboratory) build Sci-MMR, a benchmark that checks whether multimodal research agents recover the full evidence chain behind a scientific answer, and find answer accuracy runs more than 20 points ahead of evidence recovery.

Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization
Dong-Jae Koh, Young-Kyoon Suh and colleagues (Kyungpook National University) introduce DNBENCH for database normalization from 1NF to BCNF and a multi-agent method, MARS, that improves the benchmark score by 82.0% over a single prompt.

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text
DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang and JinYeong Bak (Sungkyunkwan University) show that surface noise such as typos makes LLM judges report social bias that is not in the text, so bias measured on noisy text is overestimated.

EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
Ege C. Kaya and Abolfazl Hashemi (Purdue University) analyze the update rule of EGGROLL, the low-rank evolution strategy used to fine-tune LLMs without gradients, and introduce LOO-ROLL, a leave-one-out estimator that halves estimator error at equal evaluation cost.

Thinking with Looped Flows
Ayhan Suleymanzade (EPFL) and colleagues at KAIST, Amsterdam, CMU, TU Wien and Oxford train looped models with local denoising objectives so that early recurrent updates learn to support later ones, and reach 58.8% on ARC-AGI-1.

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers
Hanhua Hong, Chenghua Lin and colleagues (Manchester, IQuest Research, Beihang and Langboat) introduce AgentActionBench for the NLPCC 2026 shared task, which grades how agents reproduce paper experiments by recording their actions rather than only checking the final repository.

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Minghao Guo, Meng Cao and colleagues (MBZUAI and USTC) build Mr.LHDR, a deep research benchmark whose questions require long chains of dependent, multimodal evidence, and find that the best system fully completes only about a third of them.

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
Makoto Fukushima, Hua-Dong Xiong and Ehsan Moradi Pari (Honda Research Institute) define the convention gap, the difference between failure rates predicted from literal messages and observed failure rates, and use Hanabi to show that human players rely on implicit conventions that AI-AI evaluation does not capture.

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Pingchen Lu, Zhongxiang Dai and colleagues (CUHK-Shenzhen, Tianjin University and NUS) treat agent skill optimization as a budgeted sequential decision problem and use a contextual bandit to decide which candidate skills are worth evaluating.

Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models
Yu-Chung Hsiao (Cisco Systems) shows that on post-2025 proprietary models, verbalized confidence is a more robust soft score for LLM-as-a-Judge than log-probabilities, which reverses the standard advice.

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
Qibai Chen (independent researcher) and Zeming Liu (Brown University) measure how well frontier LLMs resolve package version constraints across npm, PEP 440 and Cargo, and find predictable rule-specific blind spots that tool delegation removes.

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead
Haseeb Mohammed Afsar (independent researcher) probes a seeded random sample of 400 servers from the 24,135-server MCP registry and finds that fewer than half start, and that popular tool-use benchmarks contain far more duplicated tools than real MCP servers do.

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
Shenghan Zheng and Christophe Hauser (Dartmouth College), with Dawn Song (UC Berkeley) and collaborators at Amazon, BenchFlow and several universities, build BenchShield, an instrumentation layer that detects reward hacking in agent benchmarks from a formal model of each evaluation's reward-relevant events.

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
Zixiang Chen, Yuheng Lu and colleagues at Beihang University introduce JarvisGUI, a benchmark that evaluates GUI agents on workflows spanning Android, Windows and Ubuntu, where intermediate results must move between devices.

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
Ansuman Mullick and Eray Tuzun (Bilkent University) classify personal facts into a behavioral ontology and apply category-specific retention policies as deterministic functions over LLM-extracted metadata, then locate through ablation which half of the design produces which gain.