AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Inspire: Benchmarking Scientific Literature Search for Open Research Problems
Jianrong Ding (CUHK, Microsoft Research Asia intern) and colleagues at Microsoft Research Asia and CUHK introduce INSPIRE, a benchmark where agents search an open, date-gated corpus for prior work that later solved a redacted research problem, scored separately on exposure, selection and ranking.

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
Michael Hardy, Anka Reuel, Mykel Kochenderfer and Sanmi Koyejo at Stanford (with UIUC) build a Bayesian variance-decomposition framework for sparse agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index.

DAYJOB: A Benchmark for Long-Horizon Professional Work
Stephanie Finley, Liudas Panavas, Sushant Mehta, Edwin Chen and colleagues at Surge AI release DAYJOB, 130 healthcare and finance tasks written by working professionals, each estimated at 13 to 17 hours of human work and graded all-or-nothing against an expert rubric.

Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.

Cross-Benchmark Transfer from RL on Agentic Coding Tasks
Sushant Mehta, Logan Ritchie and Edwin Chen at Surge AI post-train Kimi K2.7 Code with RL alone on 1,700 expert-built coding tasks and measure gains on six external benchmarks and on harnesses never used in training.

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Dingyuan Dai, Heli Qi, Lei Liu and a large team led from Tsinghua (Jie Tang, Juanzi Li) with CMU, Yale, Waterloo and UC collaborators introduce OSWorld-Science, a benchmark of 146 tasks in which computer-use agents must operate real scientific software and produce checkable artifacts.

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang and colleagues at Meta Superintelligence Labs introduce E2E-SWE, a benchmark of 186 tasks in which a coding agent must build a complete, installable repository from a natural-language specification and an empty workspace.

DAGent: Evaluate-then-Grow Planning for Deep Research Agents
Hanwen Liu and colleagues at New York University and NYU Shanghai introduce DAGent (NeurIPS 2026), a DAG-based deep-research system that grows its task graph a batch at a time based on confidence signals from finished nodes, instead of planning the whole graph first and repairing it after failures.

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh and colleagues at Carnegie Mellon University introduce cua-speedrun, standardized infrastructure for measuring the speed and cost of computer-use agents as well as their accuracy.

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
Zeyu Gan, Zixuan Gong and Yong Liu at the Gaoling School of AI, Renmin University of China, treat harness evolution for personal agents as a learning problem and analyze it with a preference benchmark plus approximation, generalization and optimization error bounds.

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
Yun Peng, Zihan Wu and colleagues at Fudan University and City University of Hong Kong introduce LoLBench, a benchmark that tests coding agents on the full path from a human-written enhancement proposal to an implementation in a large codebase.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Raphael Shu (OpenAgents), Yusen Zhang (Columbia), Young Min Cho (Penn) and colleagues (COLM 2026) introduce AgentWorld, a benchmark for long-horizon collaboration among 3 to 20 LLM agents with asymmetric roles in an MMORPG sandbox.

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
Weida Liang, Dawn Song and colleagues from NUS, UC Berkeley, UNC and UCSB introduce AgentXploit, a two-agent system for authorized white-box security audits of AI agent codebases, plus a benchmark of 72 reproducible vulnerabilities.

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
Hongqiang Lin (Zhejiang University) with Chao Liu, Xipeng Cao and colleagues at Alibaba Group introduce EvoPathBench, a benchmark that measures self-evolving agents at each checkpoint of their memory or skill updates instead of only at the end.

OSWorld-Pro: Process-based Evaluation for Computer Use Agents
Zhilin Wang, Yi Dong and colleagues at NVIDIA introduce OSWorld-Pro, a computer-use benchmark that scores agents on each subgoal along the way instead of only on the final file or screen state.

MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes
Andy K. Zhang and colleagues at Stanford and UC Berkeley (with Percy Liang, Dan Boneh, Dawn Song and Ion Stoica) introduce MobileCybench, a benchmark that scores agent-reported exploits by replaying them and running executable probes that check whether a specific security property was violated.

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents
Fanyu Zhao, Yinsheng Li and colleagues at Fudan University and the Qwen Business Unit of Alibaba introduce RPMem, a parametric memory for agents that compiles each session into a model-independent latent memory and maps it to LoRA weights for whichever backbone is in use.

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
Chenxu Xiong, Mu Li, Alex Smola and colleagues at Boson AI introduce MSI-Bench, a benchmark for voice agents in conversations with several speakers, such as meetings and households.

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
Yining Hua (Harvard, Agent Evaluation Science) and Levi Lian (Raycaster, Stanford) introduce WorkWorlds, an evaluation infrastructure that fixes an organization's state before any task is written, so benchmark construction cannot pre-select the evidence an agent needs.

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents
Peng Kuang, Minghao Wu and colleagues at Alibaba Token Hub (with UIUC, Northeastern and Monash) introduce BabelArena, a benchmark that ports existing English agent benchmarks into 23 languages while keeping tasks and graders executable.

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Jingjie Ning, Xueqi Li and Yibo Kong of Carnegie Mellon, with Dongting Li of Tsinghua, introduce WhatWorkedBench, which scores research agents on whether they correctly predict how component changes affect results after a limited experiment budget.

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety
Jiapeng Sun, Sirui Han, Yike Guo and colleagues at HKUST introduce PASTABench, a benchmark for whether a monitor can decide during a multi-turn agent trajectory whether to intervene, when, and on which risk (EMNLP 2026).

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
Edward Lue Chee Lip, Ivan Bercovich and colleagues audit a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record to ask what it means when no agent solves a benchmark task.