🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

Haolin Yang, Sirui Han, Yike Guo and colleagues at HKUST (with Microsoft Research, Tsinghua and the University of Macau) propose HarnessSQL, which trains SQL agents inside the same execution harness they use at deployment.

01Agents
Code Understanding is a Bottleneck for Coding Agents

Code Understanding is a Bottleneck for Coding Agents

Nishant Balepur, Kiran Tomlinson and Tobias Schnabel (Microsoft Research, with the University of Maryland) present CABRA, a synthetic benchmark that builds coding tasks as call-graph transformations to isolate which abilities drive coding-agent errors.

02Code
Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

Chen Wu, Josh Passenger and Yin Song (AWS) trace how a stateless coding agent forms, carries and abandons knowledge across ARC-AGI-3 levels by following every belief it commits to a file. Accepted at the NeurIPS 2026 CL4FMAgents workshop.

03Agents
One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

Chaoliang Yan, Yuekang Li and colleagues at UNSW present the first empirical study of conflicts between co-installed agent skills that do the same job in coding agents.

04Agents
TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

Radhika Gaonkar (Prime Intellect) introduces TRACE, a protocol that tests whether a change in an agent's verifier score reflects a change in the agent or a change in the evaluation.

05Agents
Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Yuan Tian, Bing Hu and colleagues (independent researchers with UC Berkeley, Purdue and Stanford) cross seed and self-evolved harnesses with base and LoRA-trained weights to decide when an agent should be improved through its harness and when through its weights.

06Agents
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

Xing Zhang, Peiyang He and colleagues at AWS Forward Deployed Engineering make the verifier itself the evolving object in a self-improving agent loop, building inspectable graders from small deterministic drawback detectors. Accepted at the NeurIPS 2026 workshop Who Verifies the Agents?

07Agents
IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

Jiarui Liu (CMU, internship at Meta), Wen-tau Yih, Xin Luna Dong and colleagues at Meta Reality Labs introduce IdeaScientist, a three-role agent system trained with RL to generate research proposals by transferring mechanisms from analogous problems in other fields.

08Agents
TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao and Baishakhi Ray (Columbia) with Koushik Sen and Dawn Song (UC Berkeley) introduce TestJack, an auditor that generates per-trial tests for coding-agent patches and finds that about a third of trials currently scored correct violate the task requirements.

09Evaluation
BrickBench: Evaluating Agentic Brick Design

BrickBench: Evaluating Agentic Brick Design

Peter Kulits, Jiajun Wu and colleagues at Stanford (with Max Planck and Inria's Cordelia Schmid) introduce BrickBench, a benchmark where coding agents design LEGO assemblies from text prompts that must also be physically buildable.

10Evaluation
Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Jimmy Lin and colleagues at the University of Waterloo describe Project Greenhouse, an effort to build fully open and sovereign models for agentic search on modest compute, starting with Gaggle, a pointwise decoder-only reranker pre-trained from scratch.

11Agents
Recursive Self-Improvement through Multi-Agent Self-Supervision

Recursive Self-Improvement through Multi-Agent Self-Supervision

Hyunin Lee (UC Berkeley, intern at Sakana AI), Yujin Tang, Jinglue Xu, Matei Zaharia and colleagues propose Multi-Agent Self-Supervision (MASS), a recursive self-improvement method for non-verifiable tasks where the model is its own optimizer and evaluator.

12Agents
Mental-Models for Multi-Agent Systems

Mental-Models for Multi-Agent Systems

Hanan Gani, Lulu Shao and Manmohan Chandraker (UC San Diego) equip agents with a learned latent mental model of their counterpart and use it as a decision variable for action selection. Accepted at NeurIPS 2026.

13Agents
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Xing Han Lù, Siva Reddy, Alexandre Drouin, Christopher Pal and colleagues at McGill, Mila and ServiceNow Research release AgentHorizon, a benchmark for judges that decide whether long computer-use trajectories actually completed the instruction.

14Agents
Humanize: Judgement Engineering for Agentic Coding

Humanize: Judgement Engineering for Agentic Coding

Sihao Liu, Ligeng Zhu, Song Han and colleagues at NVIDIA (with UCLA, MIT and Tsinghua) describe Humanize, an open-source builder-reviewer workflow for agentic coding in which deterministic hooks, not a model, decide when work moves between planning, implementation, review and learning.

15Agents
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

Yuyao Ge and colleagues at the Institute of Computing Technology, Chinese Academy of Sciences (with UC Merced and Tsinghua) present SkillForge, an agentic RL method in which the skill library and the policy are updated together. Accepted at NeurIPS 2026.

16Agents
Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Yupu Wang, Zhengyuan Jiang, Reachal Wang and Neil Zhenqiang Gong (Duke) introduce the package hallucination attack, in which a poisoned rule file such as AGENTS.md, CLAUDE.md or .cursorrules makes a coding agent import an attacker-controlled package instead of a legitimate dependency.

17Agents
ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

Arkajyoti Chakraborty, Andreas Stolcke and colleagues at Uniphore (with UIUC) present ToolRACER, a pipeline that coordinates user, assistant and tool emulator models to generate validated multi-turn tool-calling conversations, many with non-cooperative users.

18Agents
SpecGuard: Proving a Task Is Broken Before the Agent Cheats

SpecGuard: Proving a Task Is Broken Before the Agent Cheats

Param Biyani (MATS) and Krishnamurthy Dvijotham (Google DeepMind) present SpecGuard, which checks before an agent runs whether a coding task's description and its tests can both be satisfied, and produces a Lean 4 certificate when they cannot.

19Agents
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

20Agents
How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

Xijie Gong, Tingxu Han, Lijie Hu and colleagues at MBZUAI (with Griffith and other universities) trace how agentic LLMs decide whether to call a tool or answer directly. The paper is accepted at NeurIPS 2026.

21Agents
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

Egor Pakhomov and Erik Nijkamp (Salesforce AI Research) test what MemoryAgentBench's Conflict Resolution split measures by running its own stated rule, newest statement wins, as a zero-learning resolver. Accepted at the NeurIPS 2026 Interpreting Agent Behavior workshop.

22Agents
CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

Jixuan Chen (UC San Diego) with Jiaxin Zhang, Silvio Savarese, Chien-Sheng Wu and colleagues at Salesforce AI Research and the University of Washington introduce CoTrace, a data recipe for training terminal agents while their harness is also being improved.

23Agents
Training Advisors for LLM Agents from Task Outcomes

Training Advisors for LLM Agents from Task Outcomes

Sergei Polezhaev, Barys Liskavets, Ori Press and Alexander Golubev (Nebius AI) introduce Caddie, which trains a critic model to give natural-language advice to a frozen agent mid-task, using only whether the agent eventually succeeds as reward.

24Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026