AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
Haolin Yang, Sirui Han, Yike Guo and colleagues at HKUST (with Microsoft Research, Tsinghua and the University of Macau) propose HarnessSQL, which trains SQL agents inside the same execution harness they use at deployment.

Code Understanding is a Bottleneck for Coding Agents
Nishant Balepur, Kiran Tomlinson and Tobias Schnabel (Microsoft Research, with the University of Maryland) present CABRA, a synthetic benchmark that builds coding tasks as call-graph transformations to isolate which abilities drive coding-agent errors.

Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
Chen Wu, Josh Passenger and Yin Song (AWS) trace how a stateless coding agent forms, carries and abandons knowledge across ARC-AGI-3 levels by following every belief it commits to a file. Accepted at the NeurIPS 2026 CL4FMAgents workshop.

One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
Chaoliang Yan, Yuekang Li and colleagues at UNSW present the first empirical study of conflicts between co-installed agent skills that do the same job in coding agents.

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
Radhika Gaonkar (Prime Intellect) introduces TRACE, a protocol that tests whether a change in an agent's verifier score reflects a change in the agent or a change in the evaluation.

Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Yuan Tian, Bing Hu and colleagues (independent researchers with UC Berkeley, Purdue and Stanford) cross seed and self-evolved harnesses with base and LoRA-trained weights to decide when an agent should be improved through its harness and when through its weights.

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
Xing Zhang, Peiyang He and colleagues at AWS Forward Deployed Engineering make the verifier itself the evolving object in a self-improving agent loop, building inspectable graders from small deterministic drawback detectors. Accepted at the NeurIPS 2026 workshop Who Verifies the Agents?

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation
Jiarui Liu (CMU, internship at Meta), Wen-tau Yih, Xin Luna Dong and colleagues at Meta Reality Labs introduce IdeaScientist, a three-role agent system trained with RL to generate research proposals by transferring mechanisms from analogous problems in other fields.

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
Shuangjie Yao and Baishakhi Ray (Columbia) with Koushik Sen and Dawn Song (UC Berkeley) introduce TestJack, an auditor that generates per-trial tests for coding-agent patches and finds that about a third of trials currently scored correct violate the task requirements.

BrickBench: Evaluating Agentic Brick Design
Peter Kulits, Jiajun Wu and colleagues at Stanford (with Max Planck and Inria's Cordelia Schmid) introduce BrickBench, a benchmark where coding agents design LEGO assemblies from text prompts that must also be physically buildable.

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
Jimmy Lin and colleagues at the University of Waterloo describe Project Greenhouse, an effort to build fully open and sovereign models for agentic search on modest compute, starting with Gaggle, a pointwise decoder-only reranker pre-trained from scratch.

Recursive Self-Improvement through Multi-Agent Self-Supervision
Hyunin Lee (UC Berkeley, intern at Sakana AI), Yujin Tang, Jinglue Xu, Matei Zaharia and colleagues propose Multi-Agent Self-Supervision (MASS), a recursive self-improvement method for non-verifiable tasks where the model is its own optimizer and evaluator.

Mental-Models for Multi-Agent Systems
Hanan Gani, Lulu Shao and Manmohan Chandraker (UC San Diego) equip agents with a learned latent mental model of their counterpart and use it as a decision variable for action selection. Accepted at NeurIPS 2026.

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
Xing Han Lù, Siva Reddy, Alexandre Drouin, Christopher Pal and colleagues at McGill, Mila and ServiceNow Research release AgentHorizon, a benchmark for judges that decide whether long computer-use trajectories actually completed the instruction.

Humanize: Judgement Engineering for Agentic Coding
Sihao Liu, Ligeng Zhu, Song Han and colleagues at NVIDIA (with UCLA, MIT and Tsinghua) describe Humanize, an open-source builder-reviewer workflow for agentic coding in which deterministic hooks, not a model, decide when work moves between planning, implementation, review and learning.

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
Yuyao Ge and colleagues at the Institute of Computing Technology, Chinese Academy of Sciences (with UC Merced and Tsinghua) present SkillForge, an agentic RL method in which the skill library and the policy are updated together. Accepted at NeurIPS 2026.

Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
Yupu Wang, Zhengyuan Jiang, Reachal Wang and Neil Zhenqiang Gong (Duke) introduce the package hallucination attack, in which a poisoned rule file such as AGENTS.md, CLAUDE.md or .cursorrules makes a coding agent import an attacker-controlled package instead of a legitimate dependency.

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Arkajyoti Chakraborty, Andreas Stolcke and colleagues at Uniphore (with UIUC) present ToolRACER, a pipeline that coordinates user, assistant and tool emulator models to generate validated multi-turn tool-calling conversations, many with non-cooperative users.

SpecGuard: Proving a Task Is Broken Before the Agent Cheats
Param Biyani (MATS) and Krishnamurthy Dvijotham (Google DeepMind) present SpecGuard, which checks before an agent runs whether a coding task's description and its tests can both be satisfied, and produces a Lean 4 certificate when they cannot.

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
Xijie Gong, Tingxu Han, Lijie Hu and colleagues at MBZUAI (with Griffith and other universities) trace how agentic LLMs decide whether to call a tool or answer directly. The paper is accepted at NeurIPS 2026.

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
Egor Pakhomov and Erik Nijkamp (Salesforce AI Research) test what MemoryAgentBench's Conflict Resolution split measures by running its own stated rule, newest statement wins, as a zero-learning resolver. Accepted at the NeurIPS 2026 Interpreting Agent Behavior workshop.

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
Jixuan Chen (UC San Diego) with Jiaxin Zhang, Silvio Savarese, Chien-Sheng Wu and colleagues at Salesforce AI Research and the University of Washington introduce CoTrace, a data recipe for training terminal agents while their harness is also being improved.

Training Advisors for LLM Agents from Task Outcomes
Sergei Polezhaev, Barys Liskavets, Ori Press and Alexander Golubev (Nebius AI) introduce Caddie, which trains a critic model to give natural-language advice to a frozen agent mid-task, using only whether the agent eventually succeeds as reward.