AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
A 28-author team led by Yanzhang Ma and Zhenghan Tai treats post-deployment improvement of a financial QA system as controlled behavioral maintenance, where each recurring failure becomes a scoped skill patch that must earn deployment without causing regressions.

DeltaSelect: Affordable A/B Testing for Coding Agents
Nicholas J. Conn presents DeltaSelect, which picks a fixed set of benchmark tasks whose single-run results track full-benchmark performance within a dollar budget, for repeated baseline-versus-candidate comparisons during development.

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
Yinzhu Quan and Zefang Liu distill verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data and separate skill transfer from skill retrieval.

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
Mahsa Amani and colleagues run the first end-to-end study of agentic Web search across ChatGPT, Claude, Grok and DeepSeek, combining real user interactions with controlled API experiments on the same models.

AgentPProf: Semantic Profiler for Long Horizon AI Agents
Yusheng Zheng and colleagues adapt systems profiling to agent trajectories with a semantic operation stack, so resource use can be attributed to task intent rather than to code paths.

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
Google Cloud AI Research built ScientistTwo, a multi-agent framework that takes a problem from a human expert and runs the full discovery cycle without further intervention, from establishing baselines and screening ideas on a data subset to running its own ablations and revising the idea from them. Manuscript drafting includes a simulated peer-review and rebuttal engine. Benchmarked on problems from papers accepted at ICLR, ICML, and NeurIPS, its solutions outperform the human state-of-the-art models, and its papers score higher average ratings than the human-authored ones under automated AI reviewers.

Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
Alex Remedios, Simon Storf, Fabien Roger and John Hughes at Anthropic red-team the production blocking monitors that review each coding-agent action (Auto Mode in Claude Code, Guardian in Codex) under the assumption that the monitored agent is itself persistently misaligned.

Quantifying Overclaiming Propensity in Frontier LLM Agents
Nolan Smyth, Yorguin-Jose Mantilla-Ramos and colleagues at Tara Research introduce OverclaimBench, which defines overclaiming as a final response that contradicts the agent's own context and measures it on eight proprietary frontier models in their production CLIs plus four open-weight models.

Rethinking Multi-Agent Collaboration: When More Is Less
Yishuo Yuan and colleagues at Shanghai Jiao Tong University and Nanjing University map where multi-agent collaboration beats a single-agent harness and propose SAIGE, which models collaboration as a semantic dependency graph that grows on demand.

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
Erik Nijkamp, Anurag Koul, Egor Pakhomov and Bo Pang at Salesforce AI Research argue that a long-horizon agent must run continually without forgetting before it can learn continually, and place that capability in the harness rather than the model.

Long-horizon autoformalization of a core theorem underlying MIP* = RE
Sirui Lu, Ruixuan Deng, Yanqiao Zhu and Zhengfeng Ji present FormalFlow, which coordinates AI proving agents under human supervision, and use it to complete a machine-checked Lean 4 proof of the quantum soundness of the classical low individual-degree test, a core theorem underlying MIP* = RE.

An Empirical Study of Harness Design for Coding Agents
Run-Ze Fan and colleagues at UMass Amherst, Emory, UNC Charlotte and Zoom hold a coding harness's execution loop fixed and vary three components (planning, action space, context management) across 176 matched settings, four models, SWE-Bench Verified and Terminal-Bench 2.1.

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
Tisha Chawla and Susheem Koul at Microsoft present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays a chosen subset of them while running the rest live, turning a recorded incident into a CI regression test.

Do AI Agents Understand Computer Architecture?
Ambika Sharan, Grigory Chirkov and Soheil Abbasloo at Microsoft Research build AutoTuring, which gives the same agent the same 15-dimensional accelerator design space twice, once with named architectural knobs and simulator counters and once as anonymous variables on [0,1].

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
Agent harnesses are tuned by hand, one mechanism at a time, against whatever environment the team happens to have. NVIDIA moves that tuning into an automated research loop and keeps only the mechanisms that survive selection across many environments.

ClashBench: Conflicts Leading Agents to Seize and Harm
Yuejin Xie and colleagues at Tsinghua, Shanghai AI Lab, Fudan, HKUST and KAUST name destructive resource preemption, where an agent obtains what a task needs by terminating or degrading an incumbent task, and measure it with ClashBench.

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents
Yukun Zhang, Kemu Xu and Yishen Chen at CUHK and the University of Edinburgh separate what a harness contributes by pairing real task-specific plans against shuffled policy text matched in word count, across two Retail experiments and an Airline pilot in tau^2-bench.

Closed-World Resolution Against Tool Hallucination in LLM Agents
Laxmipriya Ganesh Iyer shows that tool-selection and tool-gating defenses cannot address calls to tools that do not exist, gives a five-class taxonomy of tool hallucination, and measures the problem across ten hosted models and the Model Context Protocol.

SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale
Md Tahmid Rahman Laskar, Xue-Yong Fu and Shashi Bhushan TN at Dialpad compare LoRA SFT, GRPO and SFT followed by GRPO for tool calling across six Qwen3 models from 0.6B to 32B, measuring both in-distribution accuracy and cross-dataset transfer.

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
Yipeng Liu and colleagues at Tsinghua University, Zhejiang University and Alibaba Cloud argue that serving systems should read progress reports from running tool calls, instead of predicting tool duration, when deciding whether an agent's KV cache stays on the GPU during a tool wait.

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
Franziska Roesner (University of Washington) and Tadayoshi Kohno (Georgetown University) adapt Thompson's trusting-trust attack to self-modifying coding agents and show that a poisoned self-evaluation benchmark can make later agent versions write vulnerable code on clean, held-out tasks.

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
Bofan Chen, Boxuan Zhang and colleagues at Zhejiang University and UESTC present EvoSkill-GUI, a training-free framework in which GUI agent skills are multi-file packages that the agent revises from execution feedback during deployment.

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Jeonghye Kim (KAIST) with Microsoft Research Montréal and Microsoft AI collaborators introduce ProgramDistill, a benchmark where coding agents must infer features from a working reference web application and implement them in an incomplete copy.

Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery
Xiangfan Wu and colleagues at Tencent Zhuque Lab model collective loss of control in multi-agent LLM systems as an epidemic of mutation, contagion and recovery, and test two parts of that model with a deployment audit and the RogueHandoff-20 benchmark.