🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

49Agents
SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Jennifer Williams, Dave Farris, Jeff Farris and Jiantao Jiao (NVIDIA and UC Berkeley) introduce SWE-Serve, 53 repository-level tasks taken from recent production changes to SGLang, to test whether coding agents can implement inference-serving features correctly.

50Evaluation
Recursive self-improvement of AI research agents

Recursive self-improvement of AI research agents

Dhruv Srikanth, Zhengyao Jiang and colleagues at Weco AI present AIDE^2, a loop in which a frontier AI research agent edits its own code, benchmarks the new versions on AI R&D tasks and keeps the edits that score best on hidden evaluations.

51Agents
MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Ruike Cao, Fanyu Zhao and colleagues at Alibaba's Qwen Applications Business Group, USTC and Fudan introduce MemCalib, a benchmark for whether a model gives each retrieved memory the right amount of influence on its answer, and MemCalib-RL to train that behavior.

52Evaluation
DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Soumil Rathi, Deshraj Yadav and Taranjeet Singh (Mem0) release DolphinBench, a memory benchmark that scores agents on actions they take in simulated apps rather than on answers to recall questions, and requires every submission to report cost and latency with accuracy.

53Agents
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Running a frontier LLM as the judge on every eval gets expensive at scale. This paper tests a cheaper setup, where a decision-only judge handles most calls and only the uncertain ones go to a frontier model.

54Evaluation
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Chen-Yu Lee, Tomas Pfister and colleagues at Google Cloud AI Research, with UNC Chapel Hill, Stanford and WashU, show that automated harness evolution overfits its training tasks and add regularization to both the edit proposer and the selector.

55Agents
Self-Organizing Agent Teams Learn to Reason Together

Self-Organizing Agent Teams Learn to Reason Together

Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

56Agents
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

Laskar, Fu and colleagues (Dialpad) test whether improving next-turn metrics under gold history predicts better autonomous multi-turn workflow execution, and find that it does not.

57Agents
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

An, Jang, Kim, Lee, Park and Lee (KRAFTON) release AgentVidBench, a multi-hop video QA benchmark that tests spatial, temporal and causal reasoning in multimodal agents and scores the solution trajectory as well as the final answer.

58Evaluation
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Yining She (Carnegie Mellon, work done at Meta) and Lei Lin (Meta) study how to re-evaluate a production analytics agent with tens of thousands of monthly users without rerunning its full benchmark each time, using 574 historical benchmark runs.

59Evaluation
GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

Xinyu Che, Jiaheng Liu and colleagues at Nanjing University build GameLogicBench, 72 gameplay-logic tasks in Godot projects whose rules are checked at every simulation tick, and use it to show where coding agents fail on runtime behavior.

60Agents
Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Xinshuai Guo, Junjie Wu and colleagues at Tencent Hunyuan and Tsinghua University propose DualViewEval, which compresses expensive agent benchmarks into small task subsets by modeling both final scores and process signals from agent trajectories.

61Agents
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

Kratika Bhagtani, Kusha Sridhar and colleagues introduce ERPBench, which evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against values in its database.

62Agents
ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

Liyang Fan, Chi Wei, Bo Li and colleagues at SIAT (Chinese Academy of Sciences), Shenzhen University and China Tower introduce ReFigBench, 1,000 real arXiv overview figures that coding agents must rebuild as editable PowerPoint slides, and use it to separate model effects from harness effects.

63Evaluation
RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

Qingnuan Han, Boli Fang and colleagues at Didi Global introduce RideWay, a ride-hailing agent benchmark with a metric that discounts successful runs for extra tool calls and extra user-facing turns.

64Agents
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

Aman Priyanshu and colleagues at Cisco Foundation AI introduce VLoc Bench, which tests whether agents can find the files affected by a known weakness class in an unfamiliar repository, and whether they can recognize that a patched version is no longer vulnerable.

65Agents
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Harsh Raj and colleagues at Scale AI treat root-cause attribution of long agent failures as a search problem and introduce Continual Search, which prompts an LLM judge over several turns to keep looking for unresolved evidence instead of settling on its first plausible diagnosis.

66Agents
VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

Yu Bai and colleagues (Zhongguancun Laboratory, Tsinghua University and China Mobile) build VRL-Bench to compare verbal trial-and-error learning methods such as Reflexion under a fixed trial budget, and propose a scheduler that splits the budget between exploiting reflections and exploring.

67Evaluation
BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

Rao and Jaggi build a measurement harness that makes the per-call input-token budget the independent variable when comparing agent memory strategies, and report budget-violation rates as a first-class outcome rather than a footnote.

68Memory
MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

Jiang, Yuan and Li build a benchmark that measures rare high-severity memory failures in long-horizon agents per risk category, on the argument that an aggregate accuracy score hides exactly the events that matter.

69Evaluation
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement

Siwei Wu, Chenghua Lin and colleagues at Beihang University, the University of Manchester and collaborating institutions propose ModularRSI, a framework for evolving agent harnesses that transfer to unseen tasks rather than overfitting the benchmark used during evolution.

70Evaluation
LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

Zhao and colleagues introduce a memory architecture that labels each write with its intended lifetime, so information meant to apply only to the current context cannot overwrite knowledge meant to persist.

71Memory
From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

Jing Jiang and colleagues present HALTER, which restores a robot workspace between rollouts by planning over a library of learned atomic reset skills, so demonstration cost scales with the library rather than with the number of terminal states.

72Robotics
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026