🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
Inspire: Benchmarking Scientific Literature Search for Open Research Problems

Inspire: Benchmarking Scientific Literature Search for Open Research Problems

Jianrong Ding (CUHK, Microsoft Research Asia intern) and colleagues at Microsoft Research Asia and CUHK introduce INSPIRE, a benchmark where agents search an open, date-gated corpus for prior work that later solved a redacted research problem, scored separately on exposure, selection and ranking.

25Evaluation
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Michael Hardy, Anka Reuel, Mykel Kochenderfer and Sanmi Koyejo at Stanford (with UIUC) build a Bayesian variance-decomposition framework for sparse agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index.

26Evaluation
DAYJOB: A Benchmark for Long-Horizon Professional Work

DAYJOB: A Benchmark for Long-Horizon Professional Work

Stephanie Finley, Liudas Panavas, Sushant Mehta, Edwin Chen and colleagues at Surge AI release DAYJOB, 130 healthcare and finance tasks written by working professionals, each estimated at 13 to 17 hours of human work and graded all-or-nothing against an expert rubric.

27Evaluation
Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.

28Evaluation
Cross-Benchmark Transfer from RL on Agentic Coding Tasks

Cross-Benchmark Transfer from RL on Agentic Coding Tasks

Sushant Mehta, Logan Ritchie and Edwin Chen at Surge AI post-train Kimi K2.7 Code with RL alone on 1,700 expert-built coding tasks and measure gains on six external benchmarks and on harnesses never used in training.

29Code
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Dingyuan Dai, Heli Qi, Lei Liu and a large team led from Tsinghua (Jie Tang, Juanzi Li) with CMU, Yale, Waterloo and UC collaborators introduce OSWorld-Science, a benchmark of 146 tasks in which computer-use agents must operate real scientific software and produce checkable artifacts.

30Agents
E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang and colleagues at Meta Superintelligence Labs introduce E2E-SWE, a benchmark of 186 tasks in which a coding agent must build a complete, installable repository from a natural-language specification and an empty workspace.

31Evaluation
DAGent: Evaluate-then-Grow Planning for Deep Research Agents

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

Hanwen Liu and colleagues at New York University and NYU Shanghai introduce DAGent (NeurIPS 2026), a DAG-based deep-research system that grows its task graph a batch at a time based on confidence signals from finished nodes, instead of planning the whole graph first and repairing it after failures.

32Agents
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh and colleagues at Carnegie Mellon University introduce cua-speedrun, standardized infrastructure for measuring the speed and cost of computer-use agents as well as their accuracy.

33Evaluation
Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

Zeyu Gan, Zixuan Gong and Yong Liu at the Gaoling School of AI, Renmin University of China, treat harness evolution for personal agents as a learning problem and analyze it with a preference benchmark plus approximation, generalization and optimization error bounds.

34Agents
LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

Yun Peng, Zihan Wu and colleagues at Fudan University and City University of Hong Kong introduce LoLBench, a benchmark that tests coding agents on the full path from a human-written enhancement proposal to an implementation in a large codebase.

35Agents
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Raphael Shu (OpenAgents), Yusen Zhang (Columbia), Young Min Cho (Penn) and colleagues (COLM 2026) introduce AgentWorld, a benchmark for long-horizon collaboration among 3 to 20 LLM agents with asymmetric roles in an MMORPG sandbox.

36Agents
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.

37Agents
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

Weida Liang, Dawn Song and colleagues from NUS, UC Berkeley, UNC and UCSB introduce AgentXploit, a two-agent system for authorized white-box security audits of AI agent codebases, plus a benchmark of 72 reproducible vulnerabilities.

38Agents
Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Hongqiang Lin (Zhejiang University) with Chao Liu, Xipeng Cao and colleagues at Alibaba Group introduce EvoPathBench, a benchmark that measures self-evolving agents at each checkpoint of their memory or skill updates instead of only at the end.

39Evaluation
OSWorld-Pro: Process-based Evaluation for Computer Use Agents

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Zhilin Wang, Yi Dong and colleagues at NVIDIA introduce OSWorld-Pro, a computer-use benchmark that scores agents on each subgoal along the way instead of only on the final file or screen state.

40Agents
MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

Andy K. Zhang and colleagues at Stanford and UC Berkeley (with Percy Liang, Dan Boneh, Dawn Song and Ion Stoica) introduce MobileCybench, a benchmark that scores agent-reported exploits by replaying them and running executable probes that check whether a specific security property was violated.

41Evaluation
RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

Fanyu Zhao, Yinsheng Li and colleagues at Fudan University and the Qwen Business Unit of Alibaba introduce RPMem, a parametric memory for agents that compiles each session into a model-independent latent memory and maps it to LoRA weights for whichever backbone is in use.

42Memory
MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

Chenxu Xiong, Mu Li, Alex Smola and colleagues at Boson AI introduce MSI-Bench, a benchmark for voice agents in conversations with several speakers, such as meetings and households.

43Agents
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

Yining Hua (Harvard, Agent Evaluation Science) and Levi Lian (Raycaster, Stanford) introduce WorkWorlds, an evaluation infrastructure that fixes an organization's state before any task is written, so benchmark construction cannot pre-select the evidence an agent needs.

44Evaluation
BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

Peng Kuang, Minghao Wu and colleagues at Alibaba Token Hub (with UIUC, Northeastern and Monash) introduce BabelArena, a benchmark that ports existing English agent benchmarks into 23 languages while keeping tasks and graders executable.

45Evaluation
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

Jingjie Ning, Xueqi Li and Yibo Kong of Carnegie Mellon, with Dongting Li of Tsinghua, introduce WhatWorkedBench, which scores research agents on whether they correctly predict how component changes affect results after a limited experiment budget.

46Agents
PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

Jiapeng Sun, Sirui Han, Yike Guo and colleagues at HKUST introduce PASTABench, a benchmark for whether a monitor can decide during a multi-turn agent trajectory whether to intervene, when, and on which risk (EMNLP 2026).

47Agents
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

Edward Lue Chee Lip, Ivan Bercovich and colleagues audit a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record to ask what it means when no agent solves a benchmark task.

48Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026