AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
Chen Wu, Josh Passenger and Yin Song (AWS) trace how a stateless coding agent forms, carries and abandons knowledge across ARC-AGI-3 levels by following every belief it commits to a file. Accepted at the NeurIPS 2026 CL4FMAgents workshop.

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Xiaomi's LLM-Core team reports MiMo-V2.6, an omni-modal MoE family (Pro at 1.02T total / 42B active, Flash at 310B / 15B active) trained by scaling RL compute across batch size, environments and grader compute.

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
Jimmy Lin and colleagues at the University of Waterloo describe Project Greenhouse, an effort to build fully open and sovereign models for agentic search on modest compute, starting with Gaggle, a pointwise decoder-only reranker pre-trained from scratch.

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
Tan Yu, Alexander Bukharin, Bryan Catanzaro, Jonathan Cohen, Jiantao Jiao and colleagues at NVIDIA (with the University of Minnesota and UC Berkeley) propose ways to predict which base checkpoint will become the strongest coding agent after agentic post-training.

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.

Harness-Aware Distillation for Small Language Model Agents
Moonseok Choi, Taehong Moon, Giung Nam and Juho Lee (KAIST AI and an independent researcher) propose Harness-Aware Distillation (HAD), which distills a harness-equipped agent by teaching the student the decisions the teacher makes differently because of the harness.

Trained Agentic Context Management
Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
Ziyang Xu, Tieyong Zeng and colleagues at CUHK and the Chinese Academy of Sciences test whether populations of automatically generated LLM harnesses add real specialization, or whether their extra coverage is what rerunning a single program would give anyway.

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
Youling Huang, Lin Lin and colleagues from Kuaishou with DUT, XJTU, Tsinghua and other universities show that on-policy distillation helps agentic RL only while the teacher is ahead of the student, and propose GATS, which scales the distillation term by the measured teacher-student gap and drops it at crossover.

Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
Tong Zhang, Yihao Liu, Jiahua Bao and colleagues at Alibaba's Qwen Large Model Application Team, with Peking University and other universities, show that token-level teacher-student gaps are a poor guide to where on-policy distillation helps multi-turn agents, and propose OG-OPD, which reweights supervision using the student's actual task outcomes.

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation
Jingtan Wang and Bryan Kian Hsiang Low (NUS) with Sirajul Salekin, Young mok Jung, Javier Movellan and Manjot Bilkhu at Apple propose RISED, which uses LLM-judge rubric tags on rollouts to select training data and to supervise the policy when training one agent across several environments.

Sharpening Tax in Post-Training
Changdae Oh, Qi Zeng, Qi Qi and colleagues at Meta Superintelligence Labs, with Sharon Li (UW-Madison) and Azalia Mirhoseini (Stanford), test whether RL post-training narrows what agentic models can solve, and find that base models with a light harness often cover more tasks than their post-trained versions when given enough samples.

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Yinghui He (Princeton, NVIDIA), Jan Kautz, Ali Hatamizadeh and colleagues at NVIDIA, Princeton and UMD introduce PivotOPD, an on-policy distillation method that trains multi-turn agents both to avoid the single action that derails a rollout and to recover after making it.

SecureVibe: Making Vibe Coding More Secure
Danqing Wang (CMU), Baolin Peng, Zhepei Wei, Isadora White and colleagues at Microsoft Research with CMU and UVA introduce SecureVibe, a training recipe that targets the planning and testing behaviors that coding agents skip when they produce functionally correct but insecure code.

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Kunlun Zhu, Cheng Qian, Beibin Li, Heng Ji and colleagues at Apodex release the Agent Error Dataset (AED), 50,228 error-diagnosis pairs mined from failed agent rollouts, with a pipeline that turns failures into training data for diagnosis and recovery.

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
Bo Mao, Tao Gui, Xipeng Qiu and colleagues at East China Normal University, Fudan and Shanghai Innovation Institute introduce WEFT, which scales tool-use post-training by evolving the whole interaction system (environment, task, harness and evaluator) rather than only the environments.

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents
Xingyu Su and colleagues at AWS AI, Amazon (with Texas A&M) show that on-policy self-distillation with privileged information hurts multi-turn agents, and propose Privileged Self-Practice (PSP), which uses the privileged information only to help sample successful rollouts.

Pistis Technical Report
The Pistis team at ByteDance introduces 27B and 9B multimodal models on Qwen3.6 and Qwen3.5, trained with interleaved on-policy distillation and RL, plus an automatic harness optimizer.

Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents
Yan Zhang, Yu Zhou and colleagues at the Institute of Information Engineering (CAS), UCAS, Tencent and Tsinghua present GUI-SD-v2, which extends on-policy self-distillation from GUI grounding to multi-turn GUI interaction.

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
Pirzada Suhail, Menglin Xia and colleagues at Microsoft Research and M365 propose Pseudo Self-Distillation (PSD), which trains small Qwen3 models to run a multi-stage memory-construction pipeline that normally needs GPT-4.1-mini, using only the oracle's text outputs.

Harness-Zero: Harness Distillation via Agent-as-Harness
A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
Siyuan Liu, Yixin Cao and colleagues at Fudan University and the Meituan LongCat Team introduce DENSE, which turns an agent's own execution traces into structured feedback for a second attempt without needing outcome labels, verifiers or expert annotation.