🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

Bojian Xiong and a 14-author team score a DeepSearch run across its whole pipeline rather than only its final answer, and release a benchmark for reward models in that setting.

73Evaluation
From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

Jing Jiang and colleagues present HALTER, which restores a robot workspace between rollouts by planning over a library of learned atomic reset skills, so demonstration cost scales with the library rather than with the number of terminal states.

74Robotics
What Does Privileged Information Add to On-Policy Self-Distillation?

What Does Privileged Information Add to On-Policy Self-Distillation?

XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang and Tat-Seng Chua build a benchmark that holds the problem fixed while varying what the teacher sees, and find the privileged information adds much less than distillation itself.

75Training
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

Pyrros Koussios and colleagues introduce PetriBench, which evaluates LLM reasoning over dynamic state spaces using Petri nets, a formalism for concurrent and distributed systems, with exact ground truth and procedural generation.

76Reasoning
Self Improvement via Fast Tree-search

Self Improvement via Fast Tree-search

Coding agents that rewrite their own implementation can improve on benchmarks, but prior methods such as the Darwin Gödel Machine (DGM) are expensive to run. Researchers from MIT and Sakana AI trace most of that cost to one step and make it cheaper.

77Evaluation
For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

Alexander Shirnin and Aleksey Kudelya build a cooperative signalling game in which a Sender describes two words and an isolated Receiver, sharing only pretraining and task instructions, must identify the hidden target.

78Evaluation
DeltaSelect: Affordable A/B Testing for Coding Agents

DeltaSelect: Affordable A/B Testing for Coding Agents

Nicholas J. Conn presents DeltaSelect, which picks a fixed set of benchmark tasks whose single-run results track full-benchmark performance within a dollar budget, for repeated baseline-versus-candidate comparisons during development.

79Evaluation
SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

Mengxiao Wang and Nitesh Saxena present FARSIGHT, a scheme-level evaluation of financial LLM trading agents on robustness under market turbulence and security against three attack classes, and apply it to 15 academic schemes.

80Evaluation
Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

Xuan Liu (Shanghai Jiao Tong University) and Jingbin Qian (Rice University) introduce checkpoint handoff, which clones a state one released checkpoint reached and hands it to another, splitting an agentic RL endpoint gain into REACH and SOLVE.

81Evaluation
Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

Alex Remedios, Simon Storf, Fabien Roger and John Hughes at Anthropic red-team the production blocking monitors that review each coding-agent action (Auto Mode in Claude Code, Guardian in Codex) under the assumption that the monitored agent is itself persistently misaligned.

82Agents
Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

Leon Bergen, Usha Bhalla, Andrew Lee and colleagues at Goodfire show that difference-of-means (DoM) activation vectors identify reward hacking in Kimi K3, GLM 5.2 and Qwen 3.8 Max during coding evaluations, and that these probes cost almost nothing to run compared with LLM monitors.

83Evaluation
Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Franziska Roesner (University of Washington) and Tadayoshi Kohno (Georgetown University) adapt Thompson's trusting-trust attack to self-modifying coding agents and show that a poisoned self-evaluation benchmark can make later agent versions write vulnerable code on clean, held-out tasks.

84Evaluation
AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

Li Chen (harnets.ai) presents AutoTuneBench, a benchmark and measurement protocol for LLM agents that tune GPU kernels and serving engines, built after a four-day pilot of 619 model calls showed that the propose-measure-keep loop produces untrustworthy speedups.

85Evaluation
Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

Guojun Zhu and colleagues at the University of Chinese Academy of Sciences and the National University of Singapore introduce CHASE, which prevents automatic harness optimization from producing harnesses whose benchmark gains depend on shortcuts in the benchmark protocol rather than on the tasks.

86Evaluation
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Sadia Asif and Mohammad Mohammadi Amiri (RPI) with Prasanna Sattigeri and colleagues (IBM Research) introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents.

87Safety
The AI-Enabled Scientific Frontier

The AI-Enabled Scientific Frontier

Gabriel Manso, Emma Fu and Neil Thompson (MIT FutureTech) assemble 2,507 head-to-head comparisons between AI and other analysis methods across 27 disciplines from 2000 to early 2025.

88Evaluation
Skill-based Agentic Evaluation for Real-time Data Science Tasks

Skill-based Agentic Evaluation for Real-time Data Science Tasks

Storing a fixed reference answer for every eval case breaks when the underlying data changes daily, so Adobe researchers write each reference answer as a Python function that runs against the live system at evaluation time. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall regardless of output format, raising agreement with expert labels from an MCC of 0.331 to 0.427 while cutting token cost per case by 16%. A judge given no ground truth scored an MCC of -0.379, which is worse than chance.

89Agents
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

Bowen Qin and colleagues (NUS, PKU, CASIA, JD.com) introduce ImpossibleRubrics to test whether LLM-generated rubrics reward honest answers over adversarial answers when the only honest response is to acknowledge the task cannot be done.

90Evaluation
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

Zihan Dong, Guohao Li, Kaixin Li and colleagues (Georgia Tech, CAMEL-AI, NUS) introduce CADWorld, a long-horizon computer-use benchmark in FreeCAD graded by executable checks on the saved design files.

91Evaluation
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Fengshuo Liu (Imperial College London) and colleagues audit 254 published SWE-bench submissions without running any model and find that the top of the Verified leaderboard cannot be ordered from the published verdicts.

92Code
Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

NVIDIA compared eight strategies for choosing which models go into a multi-agent system, based on size, accuracy, answer diversity, and error diversity, across routing, majority vote, and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy while achieved accuracy often fell below the single best model in the pool, and using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, so measure what another model adds before putting it in the router.

93Agents
Verifiable Social Reasoning for LLM Assistants

Verifiable Social Reasoning for LLM Assistants

People ask assistants for social advice constantly, and the assistant only hears the user's version of events, which makes it hard to check whether it read the situation correctly. Google Research builds that ground truth by simulation, with a target agent holding a hidden motive while a user agent relays events to the assistant, which then has to infer the motive. Across 24k human annotations validating the simulations and 12 LLMs tested, biased framing from the user shifted the assistant's answer, and longer conversations with room for clarifying questions did not reliably help.

94Evaluation
The Router Within: Eliciting Native Skill Routing from a Frozen LLM

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Ruishuo Chen and colleagues at Tsinghua University show that a frozen agent LLM already carries the signal for choosing which skill to load, and build Gavel, a router that reads it out with two trained linear maps and no skill text in the context.

95Agents
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Arham Sethi and colleagues at Spark AI Research and Apta AI build a 1,024-item benchmark that forces a tool call and guarantees an unusable payload, and measure how often tool-augmented models then assert a value the tool never returned or invent a reason for withholding one.

96Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026