AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
Bojian Xiong and a 14-author team score a DeepSearch run across its whole pipeline rather than only its final answer, and release a benchmark for reward models in that setting.

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation
Jing Jiang and colleagues present HALTER, which restores a robot workspace between rollouts by planning over a library of learned atomic reset skills, so demonstration cost scales with the library rather than with the number of terminal states.

What Does Privileged Information Add to On-Policy Self-Distillation?
XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang and Tat-Seng Chua build a benchmark that holds the problem fixed while varying what the teacher sees, and find the privileged information adds much less than distillation itself.

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
Pyrros Koussios and colleagues introduce PetriBench, which evaluates LLM reasoning over dynamic state spaces using Petri nets, a formalism for concurrent and distributed systems, with exact ground truth and procedural generation.

Self Improvement via Fast Tree-search
Coding agents that rewrite their own implementation can improve on benchmarks, but prior methods such as the Darwin Gödel Machine (DGM) are expensive to run. Researchers from MIT and Sakana AI trace most of that cost to one step and make it cheaper.

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances
Alexander Shirnin and Aleksey Kudelya build a cooperative signalling game in which a Sender describes two words and an isolated Receiver, sharing only pretraining and task instructions, must identify the hidden target.

DeltaSelect: Affordable A/B Testing for Coding Agents
Nicholas J. Conn presents DeltaSelect, which picks a fixed set of benchmark tasks whose single-run results track full-benchmark performance within a dollar budget, for repeated baseline-versus-candidate comparisons during development.

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes
Mengxiao Wang and Nitesh Saxena present FARSIGHT, a scheme-level evaluation of financial LLM trading agents on robustness under market turbulence and security against three attack classes, and apply it to 15 academic schemes.

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs
Xuan Liu (Shanghai Jiao Tong University) and Jingbin Qian (Rice University) introduce checkpoint handoff, which clones a state one released checkpoint reached and hands it to another, splitting an agentic RL endpoint gain into REACH and SOLVE.

Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
Alex Remedios, Simon Storf, Fabien Roger and John Hughes at Anthropic red-team the production blocking monitors that review each coding-agent action (Auto Mode in Claude Code, Guardian in Codex) under the assumption that the monitored agent is itself persistently misaligned.

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
Leon Bergen, Usha Bhalla, Andrew Lee and colleagues at Goodfire show that difference-of-means (DoM) activation vectors identify reward hacking in Kimi K3, GLM 5.2 and Qwen 3.8 Max during coding evaluations, and that these probes cost almost nothing to run compared with LLM monitors.

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
Franziska Roesner (University of Washington) and Tadayoshi Kohno (Georgetown University) adapt Thompson's trusting-trust attack to self-modifying coding agents and show that a poisoned self-evaluation benchmark can make later agent versions write vulnerable code on clean, held-out tasks.

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
Li Chen (harnets.ai) presents AutoTuneBench, a benchmark and measurement protocol for LLM agents that tune GPU kernels and serving engines, built after a four-day pilot of 619 model calls showed that the propose-measure-keep loop produces untrustworthy speedups.

Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
Guojun Zhu and colleagues at the University of Chinese Academy of Sciences and the National University of Singapore introduce CHASE, which prevents automatic harness optimization from producing harnesses whose benchmark gains depend on shortcuts in the benchmark protocol rather than on the tasks.

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
Sadia Asif and Mohammad Mohammadi Amiri (RPI) with Prasanna Sattigeri and colleagues (IBM Research) introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents.

The AI-Enabled Scientific Frontier
Gabriel Manso, Emma Fu and Neil Thompson (MIT FutureTech) assemble 2,507 head-to-head comparisons between AI and other analysis methods across 27 disciplines from 2000 to early 2025.

Skill-based Agentic Evaluation for Real-time Data Science Tasks
Storing a fixed reference answer for every eval case breaks when the underlying data changes daily, so Adobe researchers write each reference answer as a Python function that runs against the live system at evaluation time. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall regardless of output format, raising agreement with expert labels from an MCC of 0.331 to 0.427 while cutting token cost per case by 16%. A judge given no ground truth scored an MCC of -0.379, which is worse than chance.

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
Bowen Qin and colleagues (NUS, PKU, CASIA, JD.com) introduce ImpossibleRubrics to test whether LLM-generated rubrics reward honest answers over adversarial answers when the only honest response is to acknowledge the task cannot be done.

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
Zihan Dong, Guohao Li, Kaixin Li and colleagues (Georgia Tech, CAMEL-AI, NUS) introduce CADWorld, a long-horizon computer-use benchmark in FreeCAD graded by executable checks on the saved design files.

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
Fengshuo Liu (Imperial College London) and colleagues audit 254 published SWE-bench submissions without running any model and find that the top of the Verified leaderboard cannot be ordered from the published verdicts.

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems
NVIDIA compared eight strategies for choosing which models go into a multi-agent system, based on size, accuracy, answer diversity, and error diversity, across routing, majority vote, and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy while achieved accuracy often fell below the single best model in the pool, and using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, so measure what another model adds before putting it in the router.

Verifiable Social Reasoning for LLM Assistants
People ask assistants for social advice constantly, and the assistant only hears the user's version of events, which makes it hard to check whether it read the situation correctly. Google Research builds that ground truth by simulation, with a target agent holding a hidden motive while a user agent relays events to the assistant, which then has to infer the motive. Across 24k human annotations validating the simulations and 12 LLMs tested, biased framing from the user shifted the assistant's answer, and longer conversations with room for clarifying questions did not reliably help.

The Router Within: Eliciting Native Skill Routing from a Frozen LLM
Ruishuo Chen and colleagues at Tsinghua University show that a frozen agent LLM already carries the signal for choosing which skill to load, and build Gavel, a router that reads it out with two trained linear maps and no skill text in the context.

Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
Arham Sethi and colleagues at Spark AI Research and Apta AI build a 1,024-item benchmark that forces a tool call and guarantees an unusable payload, and measure how often tool-augmented models then assert a value the tool never returned or invent a reason for withholding one.