AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
Zhuo Chen, Kewei Tu and colleagues at ShanghaiTech University represent multi-turn agent trajectories as round-level dependency DAGs and remove rounds the final answer does not depend on before fine-tuning.

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Yuan, Kang, Liu, Choi, Iyer, Jiang and Jaques (University of Washington and Stanford) propose MoDA, an online RL post-training method that counters alignment-induced mode collapse by conditioning one shared policy on numbered roles that are rewarded for producing outputs distinct from each other.

Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation
Zhiyu Gui and colleagues speed up multi-turn agentic on-policy distillation with STRIDE, which stops student rollouts once teacher support collapses and restarts generation from cached good prefixes.

Coaching Qwen3 Coder 30B to Think Like a CodeClash Arena Agent
Ivy Ning Zhang (Stanford) post-trains Qwen3-Coder-30B on stronger agents' CodeClash trajectories to improve its multi-round arena play.

What Does Privileged Information Add to On-Policy Self-Distillation?
XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang and Tat-Seng Chua build a benchmark that holds the problem fixed while varying what the teacher sees, and find the privileged information adds much less than distillation itself.

On-Demand Attention: Language Models Know When to Recall
Haibo Feng and colleagues show that a pretrained model's decoding states already predict whether a global attention read will help, and use that signal to invoke global attention selectively.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Yan Yu and colleagues find that a privileged teacher is not always reliable and that teacher supervision helps only at certain training stages, and propose RetireOPD, where the student drops the teacher on its own.

Stress-testing Alignment Midtraining
Sid Baines and colleagues test the assumptions behind alignment midtraining, which continues pretraining on alignment-relevant documents to encourage generalization, at up to 110B parameters and 1B midtraining tokens.

AutoData: Agentic Search for Pre-training Data Selection
Yan Meng and colleagues frame pretraining data selection as heuristic engineering over per-document features and introduce AutoData, an agent that searches directly over executable selection algorithms.

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Jeonghye Kim (KAIST) with Microsoft Research Montréal and Microsoft AI collaborators introduce ProgramDistill, a benchmark where coding agents must infer features from a working reference web application and implement them in an incomplete copy.

Large Language Models Develop Belief State Geometry In-Context
Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Adam Shai, Xavier Poncini and colleagues (Simplex, Astera Institute) test a computational-mechanics prediction: an LLM predicting hidden-Markov-model data in context should represent the belief state over hidden states.

Verbalizing Subliminal Learning Effects Using Text Optimization
Nathan Hu, Sanmi Koyejo and Christopher Potts (Stanford) detect subliminal learning, where distillation data carries a teacher trait that is not legible in the data, by recovering the trait as a readable prompt.

On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models
William L. Tong, Eran Malach, Emmanuel Abbe, Cengiz Pehlevan and colleagues (Apple, Harvard) show that the gating mechanism in state space models drives both their weak in-context retrieval and their better length generalization.

Data-free On-policy Distillation
Gengsheng Li and colleagues at the Institute of Automation, Chinese Academy of Sciences and Tencent find that on-policy distillation barely depends on its training data, and propose Data-free On-policy Distillation (DF-OPD), in which the teacher writes its own training questions.

Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
Arun Jose and Julian Stastny (Redwood Research) test whether synthetic document finetuning (SDF) during midtraining can inoculate a model against the broad misalignment that follows from learning to reward hack, and find that it changes what the model says without preventing the misalignment.

Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley and colleagues at Harvard Medical School fine-tune PrecepTron, a 32B judge trained with LoRA on a small number of physician-scored examples, and release GRAND-ROUNDS, a benchmark of physician scores.

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
Byte-level language models drop the tokenizer and read raw bytes, which removes a preprocessing step that no one likes but also costs accuracy at small scale. Meta studies what happens as compute grows, distilling 1B byte students from token teachers on up to 1 trillion bytes, and the ordering flips.

Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
Anqi Peter Li (Substrate Labs) and Kaden Kim (UC Berkeley) introduce the fork ledger, which measures whether an individual world-model update helped by running matched update and hold branches from the same point of a deployment stream.

The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems
Yangze Liu and Zhongyi Han (Shandong University) test whether a dominant model in an oligopoly speeds up or steers model collapse when many models retrain on a shared pool, and find that market concentration changes neither the pace nor the destination much.

EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
Ege C. Kaya and Abolfazl Hashemi (Purdue University) analyze the update rule of EGGROLL, the low-rank evolution strategy used to fine-tune LLMs without gradients, and introduce LOO-ROLL, a leave-one-out estimator that halves estimator error at equal evaluation cost.

Why Does Post-Training Quantization Work?
Yuxiang Chen, Michael Beyer, Jun Zhu and Jianfei Chen (Tsinghua University and Bosch AI Research) explain why post-training quantization of pretrained LLMs keeps accuracy even though the error from each quantized weight should compound across layers.

Legible Failures: Detecting and Repairing In-Context Binding Errors
Manas Ravulapalli, Samrath Chadha and Abhinav Hari (Efficient Computation Inc.) show that when LLMs give a wrong in-context binding, a linear probe can often read the correct binding from the hidden state, and steering toward it repairs the answer.

The information geometry of large language models is shared, learned, and controllable
Dario Picozzi (University College London) studies the Fisher-Rao geometry of next-token probabilities and shows that it is shared across transformer, state-space and recurrent language models, that it tracks what the model learns, and that it gives a principled way to make local edits with minimal side effects.

Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Rongcan Pei, Yu Meng and colleagues (University of Virginia) replace on-policy self-distillation's imitation of privileged solutions with Negative Self-Distillation, which pushes the model away from a self-generated flawed reasoner and needs no ground-truth answers.