AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
Hexuan Deng, Tianwen Jiang, Jihong Zhang and colleagues at Tencent Hy AI Data (with Beijing Zhongguancun Academy) introduce SWE-Journey, a benchmark that tests coding assistants on long development tasks with simulated users of different skill levels.

BrickBench: Evaluating Agentic Brick Design
Peter Kulits, Jiajun Wu and colleagues at Stanford (with Max Planck and Inria's Cordelia Schmid) introduce BrickBench, a benchmark where coding agents design LEGO assemblies from text prompts that must also be physically buildable.

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
Shuangjie Yao and Baishakhi Ray (Columbia) with Koushik Sen and Dawn Song (UC Berkeley) introduce TestJack, an auditor that generates per-trial tests for coding-agent patches and finds that about a third of trials currently scored correct violate the task requirements.

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
Xing Han Lù, Siva Reddy, Alexandre Drouin, Christopher Pal and colleagues at McGill, Mila and ServiceNow Research release AgentHorizon, a benchmark for judges that decide whether long computer-use trajectories actually completed the instruction.

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
Radhika Gaonkar (Prime Intellect) introduces TRACE, a protocol that tests whether a change in an agent's verifier score reflects a change in the agent or a change in the evaluation.

The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Litao Hu (Meta) and Yutong Tang (Microsoft) treat the keep-if-better step of self-improving LLM systems as selection under measurement noise and measure how much reported gains overstate held-out gains.

Coding-Agent Benchmarks Should Match Their Users' Task Flows
Igor Slinko, Yaroslav Golubev and Sergey Titov (JetBrains Research) compare 4,782 real coding-agent sessions from JetBrains IDEs with issue-derived benchmarks and propose SWE-TaskFlow to reshape benchmarks toward a measured interaction pattern.

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
Xingang Guo, Jing Gu, Jared Lichtarge and colleagues at Scale AI (with Elorian) introduce Humanity's Sixth Sense (HSS), a benchmark for the intuitive visual reasoning people perform at a glance.

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Arkajyoti Chakraborty, Andreas Stolcke and colleagues at Uniphore (with UIUC) present ToolRACER, a pipeline that coordinates user, assistant and tool emulator models to generate validated multi-turn tool-calling conversations, many with non-cooperative users.

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan, Shuowei Jin and Beidi Chen (Carnegie Mellon, Infini-AI Lab) introduce ServeLearnBench, a benchmark for agents that must infer and revise hidden environment policies from serving experience.

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas and Cozmin Ududec at the UK AI Security Institute release Transect, an open-source package built on Inspect Scout for analysing very long agent evaluation transcripts in a reproducible way.

VERA: Scaling Verifiable Environments for Agentic co-Evolution
Junqi Liu, Yucheng Tang, Daguang Xu and colleagues at NVIDIA, with UC Santa Cruz, UIUC, NUS and Tsinghua, introduce VERA, which turns benchmark trajectories into 9,000+ verifiable sandboxes and uses them to co-evolve a model and its agent harness.

Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents
Hai Dang Truong, Rayner Goh and Yintong Huo (Singapore Management University) with Thanh Le-Cong (SUTD) introduce SWE-CC, a benchmark that checks whether coding agents follow a repository's contribution policies, not only whether their patches pass tests.

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness
Yun-Yun Tsai (Columbia, during an internship at Meta) with Yuning Mao and colleagues at Meta Superintelligence Labs introduce WebUIProof, a benchmark that scores generated web interfaces by having a UI agent execute interaction tests in a headless browser.

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Zhuowen Liu at the Japan Advanced Institute of Science and Technology re-evaluates fifteen prompt-injection detectors and two task-aware judges by replaying the ground-truth tool calls of AgentDojo and tau-bench, and finds that public benchmark scores do not predict detector behavior inside agents.

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
Eray Turkel and colleagues at Roblox release OpenGameEval, a benchmark that runs language-model agents inside reproducible Roblox Studio sessions and separates observation tools from editing tools so exploration behavior can be measured directly.

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Zongxia Li, Yucheng Shi, Zhongzhi Li and colleagues at Tencent HY LLM Frontier, with the University of Maryland and others, propose Recursive Self-Rewrite (RSR), which collects successful solutions under several specialized harnesses and rewrites them into training trajectories for one general harness.

Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
Rohith Reddy Bellibatlu, Zichong Wang and Wenbin Zhang at Florida International University audit tool-using agent benchmarks by treating each tool's documented interface as an executable contract and checking whether the implementation, and the grader that trusts it, actually do what the interface says.

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain (Sarvam AI) and Armaan Sandhu (UMass Amherst) introduce UserProxyBench, which scores the simulated user in tau-bench-style agent benchmarks on whether it followed its private instructions, separately from whether the agent succeeded.

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
Mukul Singh and colleagues at Microsoft introduce EmailBench, a self-contained benchmark of 206 enterprise email and productivity scenarios built on a typed email API and a synthetic Enron-style corpus, and show that agents complete most of their tool calls correctly while failing most tasks.

When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
Ke Wang (Cambridge, Georgia Tech), Zijie Zhao (MIT) and colleagues formalize Improvement Fidelity, the requirement that an update a proxy verifier scores as better is still better after it is deployed into a world that reacts to it, and propose PIVOT-KG to spend a small budget of high-fidelity evaluations where they can change the update decision.

LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
Euntae Choi, Sumin Song and Sungjoo Yoo at Seoul National University present LiteEvo, a harness-evolution loop that starts from a benchmark-neutral harness, tells its optimizer nothing about the benchmark, costs about $8 to $13 per run, and produces harnesses that also help on held-out tasks.

LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
Bingo Zhang and colleagues at Vera Praxis with Tencent, HKUST and CUHK introduce LongPuzzleBench, 114 levels of six visual puzzle games played only through GUI input, where a legal move can make a level unsolvable without any signal until several moves later.