🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Sapien: A Stateful Policy Engine for Autonomous AI Agents

Sapien: A Stateful Policy Engine for Autonomous AI Agents

Corinn Tiffany, Wen Zhang, Eugene Bagdasarian and Lillian Tsai at Google (with UMass Amherst) present Sapien, a policy engine that enforces task-specific policies on an agent's tool calls where what is allowed depends on what the agent has already done.

121Agents
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Michael Hardy, Anka Reuel, Mykel Kochenderfer and Sanmi Koyejo at Stanford (with UIUC) build a Bayesian variance-decomposition framework for sparse agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index.

122Evaluation
Self-Evolving Coding Rules for AI Coding Agents

Self-Evolving Coding Rules for AI Coding Agents

Zhengyuan Jiang, Neil Zhenqiang Gong and colleagues at Duke University introduce RuleEvolve (NeurIPS 2026), which evolves the coding-rules files that coding agents read instead of relying on hand-written ones.

123Code
Harness Learning Enables Generalizable Test-Time Adaptation

Harness Learning Enables Generalizable Test-Time Adaptation

Alvin Zhang, Xuecheng Liu, Zixuan Wang, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette and colleagues at Carnegie Mellon University train a proposer model with RL to edit an agent's executable harness from execution feedback, and show the learned revision skill transfers to tasks it never saw.

124Agents
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

125Code
Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.

126Evaluation
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Sungho Park (POSTECH, intern at Microsoft), Jue Zhang, Pengfei Gao and colleagues at Microsoft, POSTECH and KAIST introduce ActiveSaddler, which adapts the training scenarios used to drive automated harness optimization as the harness changes.

127Agents
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Yinghui He (Princeton, NVIDIA), Jan Kautz, Ali Hatamizadeh and colleagues at NVIDIA, Princeton and UMD introduce PivotOPD, an on-policy distillation method that trains multi-turn agents both to avoid the single action that derails a rollout and to recover after making it.

128Agents
Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

Yang Cai, Vineet Gupta, Aranyak Mehta, Di Wang and colleagues at Google Research present Cogentic, a multi-agent harness built on Gemini that works on open research problems in theoretical computer science and produces natural-language proofs that domain experts then verify.

129Agents
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh and colleagues at Carnegie Mellon University introduce cua-speedrun, standardized infrastructure for measuring the speed and cost of computer-use agents as well as their accuracy.

130Evaluation
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Zheng Chen, Linfeng Liu, Hong Li and Hong Yan at Meta present RankEvolve, an auto-research framework for generative ranking models that treats execution accuracy (whether each code change is implemented correctly) as the main constraint on long research runs.

131Agents
DAGent: Evaluate-then-Grow Planning for Deep Research Agents

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

Hanwen Liu and colleagues at New York University and NYU Shanghai introduce DAGent (NeurIPS 2026), a DAG-based deep-research system that grows its task graph a batch at a time based on confidence signals from finished nodes, instead of planning the whole graph first and repairing it after failures.

132Agents
AIM: Agentic Idea Management for Automated Research

AIM: Agentic Idea Management for Automated Research

Hyeong Kyu Choi, Bhavana Dalvi Mishra, Chun-Liang Li and colleagues at Google Cloud AI Research and UW-Madison introduce AIM (Agentic Idea Manager), which organizes automated research around explicit research ideas instead of directly editing solution code.

133Agents
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Kirill Brilliantov, Alejandro Hernandez-Cano and Emmanuel Abbe at EPFL and Apple test whether elaborate MLE-agent harnesses help strong models, comparing them under matched budgets with Malena, a single-session coding agent that has only read, write and bash tools.

134Agents
E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang and colleagues at Meta Superintelligence Labs introduce E2E-SWE, a benchmark of 186 tasks in which a coding agent must build a complete, installable repository from a natural-language specification and an empty workspace.

135Evaluation
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang (KAIST, intern at NVIDIA), Byung-Kwan Lee, Yu-Chiang Frank Wang and colleagues at NVIDIA introduce Mid-Harness, which spends test-time compute at the boundary between model and harness: it samples several candidate shell actions and verifies them before one is executed.

136Agents
From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs

From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs

Kunal Jha, Max Kleiman-Weiner and Natasha Jaques at the University of Washington ask whether self-improving LLM agents that each pursue their own reward can learn from one another well enough to improve the whole population, a capability they call recursive social improvement.

137Agents
Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

Jeffrey Willette, Krishna C. Puvvada and Boris Ginsburg at NVIDIA introduce Long-Transduction, a controlled diagnostic for whether a model can keep applying state-dependent operations correctly across a long generation, the basic capability long-horizon agents depend on.

138Agents
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Kunlun Zhu, Cheng Qian, Beibin Li, Heng Ji and colleagues at Apodex release the Agent Error Dataset (AED), 50,228 error-diagnosis pairs mined from failed agent rollouts, with a pipeline that turns failures into training data for diagnosis and recovery.

139Agents
Schema: Discovering Unknown Environments via Agentic Program Induction

Schema: Discovering Unknown Environments via Agentic Program Induction

Guanning Zeng, Angjoo Kanazawa, Andrea Zanette, Haiwen Feng and colleagues at UC Berkeley, Carnegie Mellon and Impossible Research introduce Schema, an agent harness in which the LLM records what it learns about an unknown environment as executable programs instead of prose notes.

140Agents
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana (Georgia Tech) with Nikos Kanakaris, Sahika Genc and colleagues at AWS AI Labs, plus CMU and WashU, introduce MILO (Meta-evolutionary Island Orchestration), an automated harness-discovery framework that evolves the search strategy along with the harness it is searching for.

141Agents
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems

Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems

Deema Alnuhait and Hao Peng (UIUC), with Gengyu Wang (Genies) and Muhammad Khalifa (NVIDIA), show that benign LLM agents with no adversarial instruction will disguise a secret to help another agent and slip it past a monitor, a behavior they call covert assistance.

142Agents
Learning from Research: Toward Lifelong Agent Harness Evolution

Learning from Research: Toward Lifelong Agent Harness Evolution

Jingbo Yang (UCSB, intern at Microsoft), Kwei-Herng Lai, Evgeniy Gabrilovich, Shiyu Chang and colleagues at Microsoft and UC Santa Barbara introduce ScholarEvolve, which uses published agent research as the source of candidate changes when evolving a harness around a fixed model.

143Agents
SecureVibe: Making Vibe Coding More Secure

SecureVibe: Making Vibe Coding More Secure

Danqing Wang (CMU), Baolin Peng, Zhepei Wei, Isadora White and colleagues at Microsoft Research with CMU and UVA introduce SecureVibe, a training recipe that targets the planning and testing behaviors that coding agents skip when they produce functionally correct but insecure code.

144Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026