AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Sapien: A Stateful Policy Engine for Autonomous AI Agents
Corinn Tiffany, Wen Zhang, Eugene Bagdasarian and Lillian Tsai at Google (with UMass Amherst) present Sapien, a policy engine that enforces task-specific policies on an agent's tool calls where what is allowed depends on what the agent has already done.

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
Michael Hardy, Anka Reuel, Mykel Kochenderfer and Sanmi Koyejo at Stanford (with UIUC) build a Bayesian variance-decomposition framework for sparse agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index.

Self-Evolving Coding Rules for AI Coding Agents
Zhengyuan Jiang, Neil Zhenqiang Gong and colleagues at Duke University introduce RuleEvolve (NeurIPS 2026), which evolves the coding-rules files that coding agents read instead of relying on hand-written ones.

Harness Learning Enables Generalizable Test-Time Adaptation
Alvin Zhang, Xuecheng Liu, Zixuan Wang, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette and colleagues at Carnegie Mellon University train a proposer model with RL to edit an agent's executable harness from execution feedback, and show the learned revision skill transfers to tasks it never saw.

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Sungho Park (POSTECH, intern at Microsoft), Jue Zhang, Pengfei Gao and colleagues at Microsoft, POSTECH and KAIST introduce ActiveSaddler, which adapts the training scenarios used to drive automated harness optimization as the harness changes.

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Yinghui He (Princeton, NVIDIA), Jan Kautz, Ali Hatamizadeh and colleagues at NVIDIA, Princeton and UMD introduce PivotOPD, an on-policy distillation method that trains multi-turn agents both to avoid the single action that derails a rollout and to recover after making it.

Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
Yang Cai, Vineet Gupta, Aranyak Mehta, Di Wang and colleagues at Google Research present Cogentic, a multi-agent harness built on Gemini that works on open research problems in theoretical computer science and produces natural-language proofs that domain experts then verify.

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh and colleagues at Carnegie Mellon University introduce cua-speedrun, standardized infrastructure for measuring the speed and cost of computer-use agents as well as their accuracy.

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
Zheng Chen, Linfeng Liu, Hong Li and Hong Yan at Meta present RankEvolve, an auto-research framework for generative ranking models that treats execution accuracy (whether each code change is implemented correctly) as the main constraint on long research runs.

DAGent: Evaluate-then-Grow Planning for Deep Research Agents
Hanwen Liu and colleagues at New York University and NYU Shanghai introduce DAGent (NeurIPS 2026), a DAG-based deep-research system that grows its task graph a batch at a time based on confidence signals from finished nodes, instead of planning the whole graph first and repairing it after failures.

AIM: Agentic Idea Management for Automated Research
Hyeong Kyu Choi, Bhavana Dalvi Mishra, Chun-Liang Li and colleagues at Google Cloud AI Research and UW-Madison introduce AIM (Agentic Idea Manager), which organizes automated research around explicit research ideas instead of directly editing solution code.

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Kirill Brilliantov, Alejandro Hernandez-Cano and Emmanuel Abbe at EPFL and Apple test whether elaborate MLE-agent harnesses help strong models, comparing them under matched budgets with Malena, a single-session coding agent that has only read, write and bash tools.

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang and colleagues at Meta Superintelligence Labs introduce E2E-SWE, a benchmark of 186 tasks in which a coding agent must build a complete, installable repository from a natural-language specification and an empty workspace.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Minki Kang (KAIST, intern at NVIDIA), Byung-Kwan Lee, Yu-Chiang Frank Wang and colleagues at NVIDIA introduce Mid-Harness, which spends test-time compute at the boundary between model and harness: it samples several candidate shell actions and verifies them before one is executed.

From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs
Kunal Jha, Max Kleiman-Weiner and Natasha Jaques at the University of Washington ask whether self-improving LLM agents that each pursue their own reward can learn from one another well enough to improve the whole population, a capability they call recursive social improvement.

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
Jeffrey Willette, Krishna C. Puvvada and Boris Ginsburg at NVIDIA introduce Long-Transduction, a controlled diagnostic for whether a model can keep applying state-dependent operations correctly across a long generation, the basic capability long-horizon agents depend on.

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Kunlun Zhu, Cheng Qian, Beibin Li, Heng Ji and colleagues at Apodex release the Agent Error Dataset (AED), 50,228 error-diagnosis pairs mined from failed agent rollouts, with a pipeline that turns failures into training data for diagnosis and recovery.

Schema: Discovering Unknown Environments via Agentic Program Induction
Guanning Zeng, Angjoo Kanazawa, Andrea Zanette, Haiwen Feng and colleagues at UC Berkeley, Carnegie Mellon and Impossible Research introduce Schema, an agent harness in which the LLM records what it learns about an unknown environment as executable programs instead of prose notes.

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana (Georgia Tech) with Nikos Kanakaris, Sahika Genc and colleagues at AWS AI Labs, plus CMU and WashU, introduce MILO (Meta-evolutionary Island Orchestration), an automated harness-discovery framework that evolves the search strategy along with the harness it is searching for.

Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait and Hao Peng (UIUC), with Gengyu Wang (Genies) and Muhammad Khalifa (NVIDIA), show that benign LLM agents with no adversarial instruction will disguise a secret to help another agent and slip it past a monitor, a behavior they call covert assistance.

Learning from Research: Toward Lifelong Agent Harness Evolution
Jingbo Yang (UCSB, intern at Microsoft), Kwei-Herng Lai, Evgeniy Gabrilovich, Shiyu Chang and colleagues at Microsoft and UC Santa Barbara introduce ScholarEvolve, which uses published agent research as the source of candidate changes when evolving a harness around a fixed model.

SecureVibe: Making Vibe Coding More Secure
Danqing Wang (CMU), Baolin Peng, Zhepei Wei, Isadora White and colleagues at Microsoft Research with CMU and UVA introduce SecureVibe, a training recipe that targets the planning and testing behaviors that coding agents skip when they produce functionally correct but insecure code.