AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
Jennifer Williams, Dave Farris, Jeff Farris and Jiantao Jiao (NVIDIA and UC Berkeley) introduce SWE-Serve, 53 repository-level tasks taken from recent production changes to SGLang, to test whether coding agents can implement inference-serving features correctly.

Recursive self-improvement of AI research agents
Dhruv Srikanth, Zhengyao Jiang and colleagues at Weco AI present AIDE^2, a loop in which a frontier AI research agent edits its own code, benchmarks the new versions on AI R&D tasks and keeps the edits that score best on hidden evaluations.

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
Ruike Cao, Fanyu Zhao and colleagues at Alibaba's Qwen Applications Business Group, USTC and Fudan introduce MemCalib, a benchmark for whether a model gives each retrieved memory the right amount of influence on its answer, and MemCalib-RL to train that behavior.

DolphinBench: Mapping the Pareto Frontier of Agent Memory
Soumil Rathi, Deshraj Yadav and Taranjeet Singh (Mem0) release DolphinBench, a memory benchmark that scores agents on actions they take in simulated apps rather than on answers to recall questions, and requires every submission to report cost and latency with accuracy.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
Running a frontier LLM as the judge on every eval gets expensive at scale. This paper tests a cheaper setup, where a decision-only judge handles most calls and only the uncertain ones go to a frontier model.

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Peng Xia, Chen-Yu Lee, Tomas Pfister and colleagues at Google Cloud AI Research, with UNC Chapel Hill, Stanford and WashU, show that automated harness evolution overfits its training tasks and add regularization to both the edit proposer and the selector.

Self-Organizing Agent Teams Learn to Reason Together
Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success
Laskar, Fu and colleagues (Dialpad) test whether improving next-turn metrics under gold history predicts better autonomous multi-turn workflow execution, and find that it does not.

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
An, Jang, Kim, Lee, Park and Lee (KRAFTON) release AgentVidBench, a multi-hop video QA benchmark that tests spatial, temporal and causal reasoning in multimodal agents and scores the solution trajectory as well as the final answer.

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
Yining She (Carnegie Mellon, work done at Meta) and Lei Lin (Meta) study how to re-evaluate a production analytics agent with tens of thousands of monthly users without rerunning its full benchmark each time, using 574 historical benchmark runs.

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
Xinyu Che, Jiaheng Liu and colleagues at Nanjing University build GameLogicBench, 72 gameplay-logic tasks in Godot projects whose rules are checked at every simulation tick, and use it to show where coding agents fail on runtime behavior.

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
Xinshuai Guo, Junjie Wu and colleagues at Tencent Hunyuan and Tsinghua University propose DualViewEval, which compresses expensive agent benchmarks into small task subsets by modeling both final scores and process signals from agent trajectories.

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
Kratika Bhagtani, Kusha Sridhar and colleagues introduce ERPBench, which evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against values in its database.

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
Liyang Fan, Chi Wei, Bo Li and colleagues at SIAT (Chinese Academy of Sciences), Shenzhen University and China Tower introduce ReFigBench, 1,000 real arXiv overview figures that coding agents must rebuild as editable PowerPoint slides, and use it to separate model effects from harness effects.

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
Qingnuan Han, Boli Fang and colleagues at Didi Global introduce RideWay, a ride-hailing agent benchmark with a metric that discounts successful runs for extra tool calls and extra user-facing turns.

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale
Aman Priyanshu and colleagues at Cisco Foundation AI introduce VLoc Bench, which tests whether agents can find the files affected by a known weakness class in an unfamiliar repository, and whether they can recognize that a patched version is no longer vulnerable.

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Harsh Raj and colleagues at Scale AI treat root-cause attribution of long agent failures as a search problem and introduce Continual Search, which prompts an LLM judge over several turns to keep looking for unresolved evidence instead of settling on its first plausible diagnosis.

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets
Yu Bai and colleagues (Zhongguancun Laboratory, Tsinghua University and China Mobile) build VRL-Bench to compare verbal trial-and-error learning methods such as Reflexion under a fixed trial budget, and propose a scheduler that splits the budget between exploiting reflections and exploring.

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
Rao and Jaggi build a measurement harness that makes the per-call input-token budget the independent variable when comparing agent memory strategies, and report budget-violation rates as a first-class outcome rather than a footnote.

MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
Jiang, Yuan and Li build a benchmark that measures rare high-severity memory failures in long-horizon agents per risk category, on the argument that an aggregate accuracy score hides exactly the events that matter.

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Siwei Wu, Chenghua Lin and colleagues at Beihang University, the University of Manchester and collaborating institutions propose ModularRSI, a framework for evolving agent harnesses that transfer to unseen tasks rather than overfitting the benchmark used during evolution.

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory
Zhao and colleagues introduce a memory architecture that labels each write with its intended lifetime, so information meant to apply only to the current context cannot overwrite knowledge meant to persist.

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation
Jing Jiang and colleagues present HALTER, which restores a robot workspace between rollouts by planning over a library of learned atomic reset skills, so demonstration cost scales with the library rather than with the number of terminal states.