🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
392 papers · TrainingClear filters →
Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

Chen Wu, Josh Passenger and Yin Song (AWS) trace how a stateless coding agent forms, carries and abandons knowledge across ARC-AGI-3 levels by following every belief it commits to a file. Accepted at the NeurIPS 2026 CL4FMAgents workshop.

01Agents
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Xiaomi's LLM-Core team reports MiMo-V2.6, an omni-modal MoE family (Pro at 1.02T total / 42B active, Flash at 310B / 15B active) trained by scaling RL compute across batch size, environments and grader compute.

02Reinforcement Learning
Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Jimmy Lin and colleagues at the University of Waterloo describe Project Greenhouse, an effort to build fully open and sovereign models for agentic search on modest compute, starting with Gaggle, a pointwise decoder-only reranker pre-trained from scratch.

03Agents
Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Tan Yu, Alexander Bukharin, Bryan Catanzaro, Jonathan Cohen, Jiantao Jiao and colleagues at NVIDIA (with the University of Minnesota and UC Berkeley) propose ways to predict which base checkpoint will become the strongest coding agent after agentic post-training.

04Agents
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.

05Agents
Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation for Small Language Model Agents

Moonseok Choi, Taehong Moon, Giung Nam and Juho Lee (KAIST AI and an independent researcher) propose Harness-Aware Distillation (HAD), which distills a harness-equipped agent by teaching the student the decisions the teacher makes differently because of the harness.

06Training
Trained Agentic Context Management

Trained Agentic Context Management

Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

07Agents
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Ziyang Xu, Tieyong Zeng and colleagues at CUHK and the Chinese Academy of Sciences test whether populations of automatically generated LLM harnesses add real specialization, or whether their extra coverage is what rerunning a single program would give anyway.

08Training
Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Youling Huang, Lin Lin and colleagues from Kuaishou with DUT, XJTU, Tsinghua and other universities show that on-policy distillation helps agentic RL only while the teacher is ahead of the student, and propose GATS, which scales the distillation term by the measured teacher-student gap and drops it at crossover.

09Agents
Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents

Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents

Tong Zhang, Yihao Liu, Jiahua Bao and colleagues at Alibaba's Qwen Large Model Application Team, with Peking University and other universities, show that token-level teacher-student gaps are a poor guide to where on-policy distillation helps multi-turn agents, and propose OG-OPD, which reweights supervision using the student's actual task outcomes.

10Agents
RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

Jingtan Wang and Bryan Kian Hsiang Low (NUS) with Sirajul Salekin, Young mok Jung, Javier Movellan and Manjot Bilkhu at Apple propose RISED, which uses LLM-judge rubric tags on rollouts to select training data and to supervise the policy when training one agent across several environments.

11Agents
Sharpening Tax in Post-Training

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi and colleagues at Meta Superintelligence Labs, with Sharon Li (UW-Madison) and Azalia Mirhoseini (Stanford), test whether RL post-training narrows what agentic models can solve, and find that base models with a light harness often cover more tasks than their post-trained versions when given enough samples.

12Training
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Yinghui He (Princeton, NVIDIA), Jan Kautz, Ali Hatamizadeh and colleagues at NVIDIA, Princeton and UMD introduce PivotOPD, an on-policy distillation method that trains multi-turn agents both to avoid the single action that derails a rollout and to recover after making it.

13Agents
SecureVibe: Making Vibe Coding More Secure

SecureVibe: Making Vibe Coding More Secure

Danqing Wang (CMU), Baolin Peng, Zhepei Wei, Isadora White and colleagues at Microsoft Research with CMU and UVA introduce SecureVibe, a training recipe that targets the planning and testing behaviors that coding agents skip when they produce functionally correct but insecure code.

14Agents
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Kunlun Zhu, Cheng Qian, Beibin Li, Heng Ji and colleagues at Apodex release the Agent Error Dataset (AED), 50,228 error-diagnosis pairs mined from failed agent rollouts, with a pipeline that turns failures into training data for diagnosis and recovery.

15Agents
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Bo Mao, Tao Gui, Xipeng Qiu and colleagues at East China Normal University, Fudan and Shanghai Innovation Institute introduce WEFT, which scales tool-use post-training by evolving the whole interaction system (environment, task, harness and evaluator) rather than only the environments.

16Agents
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

17Agents
From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

Xingyu Su and colleagues at AWS AI, Amazon (with Texas A&M) show that on-policy self-distillation with privileged information hurts multi-turn agents, and propose Privileged Self-Practice (PSP), which uses the privileged information only to help sample successful rollouts.

18Agents
Pistis Technical Report

Pistis Technical Report

The Pistis team at ByteDance introduces 27B and 9B multimodal models on Qwen3.6 and Qwen3.5, trained with interleaved on-policy distillation and RL, plus an automatic harness optimizer.

19Multimodal
Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

Yan Zhang, Yu Zhou and colleagues at the Institute of Information Engineering (CAS), UCAS, Tencent and Tsinghua present GUI-SD-v2, which extends on-policy self-distillation from GUI grounding to multi-turn GUI interaction.

20Training
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

21Agents
PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

Pirzada Suhail, Menglin Xia and colleagues at Microsoft Research and M365 propose Pseudo Self-Distillation (PSD), which trains small Qwen3 models to run a multi-stage memory-construction pipeline that normally needs GPT-4.1-mini, using only the oracle's text outputs.

22Memory
Harness-Zero: Harness Distillation via Agent-as-Harness

Harness-Zero: Harness Distillation via Agent-as-Harness

A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

23Agents
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Siyuan Liu, Yixin Cao and colleagues at Fudan University and the Meituan LongCat Team introduce DENSE, which turns an agent's own execution traces into structured feedback for a second attempt without needing outcome labels, verifiers or expert annotation.

24Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026