🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

A 28-author team led by Yanzhang Ma and Zhenghan Tai treats post-deployment improvement of a financial QA system as controlled behavioral maintenance, where each recurring failure becomes a scoped skill patch that must earn deployment without causing regressions.

361Agents
DeltaSelect: Affordable A/B Testing for Coding Agents

DeltaSelect: Affordable A/B Testing for Coding Agents

Nicholas J. Conn presents DeltaSelect, which picks a fixed set of benchmark tasks whose single-run results track full-benchmark performance within a dollar budget, for repeated baseline-versus-candidate comparisons during development.

362Evaluation
EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

Yinzhu Quan and Zefang Liu distill verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data and separate skill transfer from skill retrieval.

363Retrieval
Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

Mahsa Amani and colleagues run the first end-to-end study of agentic Web search across ChatGPT, Claude, Grok and DeepSeek, combining real user interactions with controlled API experiments on the same models.

364Agents
AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf: Semantic Profiler for Long Horizon AI Agents

Yusheng Zheng and colleagues adapt systems profiling to agent trajectories with a semantic operation stack, so resource use can be attributed to task intent rather than to code paths.

365Agents
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

Google Cloud AI Research built ScientistTwo, a multi-agent framework that takes a problem from a human expert and runs the full discovery cycle without further intervention, from establishing baselines and screening ideas on a data subset to running its own ablations and revising the idea from them. Manuscript drafting includes a simulated peer-review and rebuttal engine. Benchmarked on problems from papers accepted at ICLR, ICML, and NeurIPS, its solutions outperform the human state-of-the-art models, and its papers score higher average ratings than the human-authored ones under automated AI reviewers.

366Agents
Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

Alex Remedios, Simon Storf, Fabien Roger and John Hughes at Anthropic red-team the production blocking monitors that review each coding-agent action (Auto Mode in Claude Code, Guardian in Codex) under the assumption that the monitored agent is itself persistently misaligned.

367Agents
Quantifying Overclaiming Propensity in Frontier LLM Agents

Quantifying Overclaiming Propensity in Frontier LLM Agents

Nolan Smyth, Yorguin-Jose Mantilla-Ramos and colleagues at Tara Research introduce OverclaimBench, which defines overclaiming as a final response that contradicts the agent's own context and measures it on eight proprietary frontier models in their production CLIs plus four open-weight models.

368Agents
Rethinking Multi-Agent Collaboration: When More Is Less

Rethinking Multi-Agent Collaboration: When More Is Less

Yishuo Yuan and colleagues at Shanghai Jiao Tong University and Nanjing University map where multi-agent collaboration beats a single-agent harness and propose SAIGE, which models collaboration as a semantic dependency graph that grows on demand.

369Agents
An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

Erik Nijkamp, Anurag Koul, Egor Pakhomov and Bo Pang at Salesforce AI Research argue that a long-horizon agent must run continually without forgetting before it can learn continually, and place that capability in the harness rather than the model.

370Agents
Long-horizon autoformalization of a core theorem underlying MIP* = RE

Long-horizon autoformalization of a core theorem underlying MIP* = RE

Sirui Lu, Ruixuan Deng, Yanqiao Zhu and Zhengfeng Ji present FormalFlow, which coordinates AI proving agents under human supervision, and use it to complete a machine-checked Lean 4 proof of the quantum soundness of the classical low individual-degree test, a core theorem underlying MIP* = RE.

371Agents
An Empirical Study of Harness Design for Coding Agents

An Empirical Study of Harness Design for Coding Agents

Run-Ze Fan and colleagues at UMass Amherst, Emory, UNC Charlotte and Zoom hold a coding harness's execution loop fixed and vary three components (planning, action space, context management) across 176 matched settings, four models, SWE-Bench Verified and Terminal-Bench 2.1.

372Code
Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Tisha Chawla and Susheem Koul at Microsoft present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays a chosen subset of them while running the rest live, turning a recorded incident into a CI regression test.

373Agents
Do AI Agents Understand Computer Architecture?

Do AI Agents Understand Computer Architecture?

Ambika Sharan, Grigory Chirkov and Soheil Abbasloo at Microsoft Research build AutoTuring, which gives the same agent the same 15-dimensional accelerator design space twice, once with named architectural knobs and simulator counters and once as anonymous variables on [0,1].

374Agents
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Agent harnesses are tuned by hand, one mechanism at a time, against whatever environment the team happens to have. NVIDIA moves that tuning into an automated research loop and keeps only the mechanisms that survive selection across many environments.

375Agents
ClashBench: Conflicts Leading Agents to Seize and Harm

ClashBench: Conflicts Leading Agents to Seize and Harm

Yuejin Xie and colleagues at Tsinghua, Shanghai AI Lab, Fudan, HKUST and KAUST name destructive resource preemption, where an agent obtains what a task needs by terminating or degrading an incumbent task, and measure it with ClashBench.

376Agents
How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Yukun Zhang, Kemu Xu and Yishen Chen at CUHK and the University of Edinburgh separate what a harness contributes by pairing real task-specific plans against shuffled policy text matched in word count, across two Retail experiments and an Airline pilot in tau^2-bench.

377Agents
Closed-World Resolution Against Tool Hallucination in LLM Agents

Closed-World Resolution Against Tool Hallucination in LLM Agents

Laxmipriya Ganesh Iyer shows that tool-selection and tool-gating defenses cannot address calls to tools that do not exist, gives a five-class taxonomy of tool hallucination, and measures the problem across ten hosted models and the Model Context Protocol.

378Safety
SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

Md Tahmid Rahman Laskar, Xue-Yong Fu and Shashi Bhushan TN at Dialpad compare LoRA SFT, GRPO and SFT followed by GRPO for tool calling across six Qwen3 models from 0.6B to 32B, measuring both in-distribution accuracy and cross-dataset transfer.

379Agents
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

Yipeng Liu and colleagues at Tsinghua University, Zhejiang University and Alibaba Cloud argue that serving systems should read progress reports from running tool calls, instead of predicting tool duration, when deciding whether an agent's KV cache stays on the GPU during a tool wait.

380Agents
Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Franziska Roesner (University of Washington) and Tadayoshi Kohno (Georgetown University) adapt Thompson's trusting-trust attack to self-modifying coding agents and show that a poisoned self-evaluation benchmark can make later agent versions write vulnerable code on clean, held-out tasks.

381Evaluation
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

Bofan Chen, Boxuan Zhang and colleagues at Zhejiang University and UESTC present EvoSkill-GUI, a training-free framework in which GUI agent skills are multi-file packages that the agent revises from execution feedback during deployment.

382Agents
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

Jeonghye Kim (KAIST) with Microsoft Research Montréal and Microsoft AI collaborators introduce ProgramDistill, a benchmark where coding agents must infer features from a working reference web application and implement them in an incomplete copy.

383Training
Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

Xiangfan Wu and colleagues at Tencent Zhuque Lab model collective loss of control in multi-agent LLM systems as an epidemic of mutation, contagion and recovery, and test two parts of that model with a deployment audit and the RogueHandoff-20 benchmark.

384Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026