🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Hexuan Deng, Tianwen Jiang, Jihong Zhang and colleagues at Tencent Hy AI Data (with Beijing Zhongguancun Academy) introduce SWE-Journey, a benchmark that tests coding assistants on long development tasks with simulated users of different skill levels.

01Code
BrickBench: Evaluating Agentic Brick Design

BrickBench: Evaluating Agentic Brick Design

Peter Kulits, Jiajun Wu and colleagues at Stanford (with Max Planck and Inria's Cordelia Schmid) introduce BrickBench, a benchmark where coding agents design LEGO assemblies from text prompts that must also be physically buildable.

02Evaluation
TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao and Baishakhi Ray (Columbia) with Koushik Sen and Dawn Song (UC Berkeley) introduce TestJack, an auditor that generates per-trial tests for coding-agent patches and finds that about a third of trials currently scored correct violate the task requirements.

03Evaluation
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Xing Han Lù, Siva Reddy, Alexandre Drouin, Christopher Pal and colleagues at McGill, Mila and ServiceNow Research release AgentHorizon, a benchmark for judges that decide whether long computer-use trajectories actually completed the instruction.

04Agents
TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

Radhika Gaonkar (Prime Intellect) introduces TRACE, a protocol that tests whether a change in an agent's verifier score reflects a change in the agent or a change in the evaluation.

05Agents
The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

Litao Hu (Meta) and Yutong Tang (Microsoft) treat the keep-if-better step of self-improving LLM systems as selection under measurement noise and measure how much reported gains overstate held-out gains.

06Safety
Coding-Agent Benchmarks Should Match Their Users' Task Flows

Coding-Agent Benchmarks Should Match Their Users' Task Flows

Igor Slinko, Yaroslav Golubev and Sergey Titov (JetBrains Research) compare 4,782 real coding-agent sessions from JetBrains IDEs with issue-derived benchmarks and propose SWE-TaskFlow to reshape benchmarks toward a measured interaction pattern.

07Evaluation
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Xingang Guo, Jing Gu, Jared Lichtarge and colleagues at Scale AI (with Elorian) introduce Humanity's Sixth Sense (HSS), a benchmark for the intuitive visual reasoning people perform at a glance.

08Multimodal
ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

Arkajyoti Chakraborty, Andreas Stolcke and colleagues at Uniphore (with UIUC) present ToolRACER, a pipeline that coordinates user, assistant and tool emulator models to generate validated multi-turn tool-calling conversations, many with non-cooperative users.

09Agents
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

10Agents
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan, Shuowei Jin and Beidi Chen (Carnegie Mellon, Infini-AI Lab) introduce ServeLearnBench, a benchmark for agents that must infer and revise hidden environment policies from serving experience.

11Agents
Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas and Cozmin Ududec at the UK AI Security Institute release Transect, an open-source package built on Inspect Scout for analysing very long agent evaluation transcripts in a reproducible way.

12Agents
VERA: Scaling Verifiable Environments for Agentic co-Evolution

VERA: Scaling Verifiable Environments for Agentic co-Evolution

Junqi Liu, Yucheng Tang, Daguang Xu and colleagues at NVIDIA, with UC Santa Cruz, UIUC, NUS and Tsinghua, introduce VERA, which turns benchmark trajectories into 9,000+ verifiable sandboxes and uses them to co-evolve a model and its agent harness.

13Agents
Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

Hai Dang Truong, Rayner Goh and Yintong Huo (Singapore Management University) with Thanh Le-Cong (SUTD) introduce SWE-CC, a benchmark that checks whether coding agents follow a repository's contribution policies, not only whether their patches pass tests.

14Evaluation
WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

Yun-Yun Tsai (Columbia, during an internship at Meta) with Yuning Mao and colleagues at Meta Superintelligence Labs introduce WebUIProof, a benchmark that scores generated web interfaces by having a UI agent execute interaction tests in a headless browser.

15Evaluation
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

Zhuowen Liu at the Japan Advanced Institute of Science and Technology re-evaluates fifteen prompt-injection detectors and two task-aware judges by replaying the ground-truth tool calls of AgentDojo and tau-bench, and finds that public benchmark scores do not predict detector behavior inside agents.

16Evaluation
OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

Eray Turkel and colleagues at Roblox release OpenGameEval, a benchmark that runs language-model agents inside reproducible Roblox Studio sessions and separates observation tools from editing tools so exploration behavior can be measured directly.

17Evaluation
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Zongxia Li, Yucheng Shi, Zhongzhi Li and colleagues at Tencent HY LLM Frontier, with the University of Maryland and others, propose Recursive Self-Rewrite (RSR), which collects successful solutions under several specialized harnesses and rewrites them into training trajectories for one general harness.

18Evaluation
Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

Rohith Reddy Bellibatlu, Zichong Wang and Wenbin Zhang at Florida International University audit tool-using agent benchmarks by treating each tool's documented interface as an executable contract and checking whether the implementation, and the grader that trusts it, actually do what the interface says.

19Agents
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

Ashish Jain (Sarvam AI) and Armaan Sandhu (UMass Amherst) introduce UserProxyBench, which scores the simulated user in tau-bench-style agent benchmarks on whether it followed its private instructions, separately from whether the agent succeeded.

20Evaluation
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

Mukul Singh and colleagues at Microsoft introduce EmailBench, a self-contained benchmark of 206 enterprise email and productivity scenarios built on a typed email API and a synthetic Enron-style corpus, and show that agents complete most of their tool calls correctly while failing most tasks.

21Evaluation
When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds

When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds

Ke Wang (Cambridge, Georgia Tech), Zijie Zhao (MIT) and colleagues formalize Improvement Fidelity, the requirement that an update a proxy verifier scores as better is still better after it is deployed into a world that reacts to it, and propose PIVOT-KG to spend a small budget of high-fidelity evaluations where they can change the update decision.

22Agents
LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

Euntae Choi, Sumin Song and Sungjoo Yoo at Seoul National University present LiteEvo, a harness-evolution loop that starts from a benchmark-neutral harness, tells its optimizer nothing about the benchmark, costs about $8 to $13 per run, and produces harnesses that also help on held-out tasks.

23Evaluation
LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

Bingo Zhang and colleagues at Vera Praxis with Tencent, HKUST and CUHK introduce LongPuzzleBench, 114 levels of six visual puzzle games played only through GUI input, where a legal move can make a level unsolvable without any signal until several moves later.

24Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026