AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
FrogNano: Training a 4B Coding Agent via Online Task Synthesis
Small coding agents are usually built by distilling a frontier model's trajectories. Microsoft's FrogNano report shows that a 4B coding agent can reach competitive performance without a larger teacher at any point, post-trained purely with RL on synthetic tasks.

Miles v0.1: Production-Level Post-Training
RadixArk releases Miles v0.1, an open-source post-training system built on the slime design, covering RL, LoRA RL, on-policy distillation and SFT, with an end-to-end case study running fully asynchronous agentic RL on a 744B-A40B GLM-5.2 model over terminal-use coding tasks.

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
The NeoHorse Team releases NeoHorse-1, a family of agent-native 4B and 9B models built on an agentic post-training loop in which a router's records of predicted capability demand and selected service tier become the training data for the next round.

Optimizer Memory Schedules for Outscaling the Overtraining Axis
Katie Everett (MIT CSAIL) and Shikai Qiu (NYU) show that optimizer rankings and optimal hyperparameters change substantially as the training horizon extends, and argue that overtraining factor belongs on the axis list for optimizer evaluation.

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection
Zhinan Hou and colleagues at Tsinghua University and Meituan study what data on-policy distillation actually needs, and find that eight hard examples match a 17K-example baseline.

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
Sanyuan Chen and colleagues at FAIR at Meta present Text-AB, a 3B latent-diffusion speech model that removes the forced-alignment stage from the Audiobox line and handles dubbing and two-speaker dialogue in one system.

Improving precipitation forecasts in an AI weather model using observational data
Julian F. Schmitt and colleagues at X, The Moonshot Factory and Google, with Caltech and Stanford, fine-tune an AI weather model on observed precipitation rather than reanalysis and improve precipitation forecasts substantially.

Normalized Low-Rank Adaptation
Jiale Kang and colleagues at CUHK, Yuanshi Intelligence and Microsoft Research normalize the down-projection matrices in LoRA and report faster convergence, better stability and less forgetting at no extra cost.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Jacqueline He and colleagues at Meta AI, the University of Washington and Princeton show that standard knowledge distillation helps reasoning and hurts factual recall during mid-training, and trace the cause to teacher confidence.

Instruction Duplication as an Inference-Time Control Primitive
Victor Lavrenko at PeaceTech VC measures instruction duplication, repeating only the procedural instruction, as a black-box inference-time control across seven instruction-tuned models and 16,800 scheduled generations, and reports gains on process compliance without any change in final-answer accuracy.

FailBench: How Reliable are VLMs at Judging Robot Task Success?
Zaruhi Navasardyan, Tatul Danielyan and Hrant Davtyan at Metric AI Lab assemble 2,197 real manipulation attempts from 14 sources and find that vision-language models used as robot success detectors reach only 0.77 mean balanced accuracy, with fine-tuned detectors doing worse than general-purpose models.

CROCODIL: Cross-Model Code Editing with LLMs
Linghan Zhong, Aditya Thimmaiah, Milos Gligoric and Junyi Jessy Li at UT Austin with Cisco Research show that a model editing code originally written by a different model makes more and larger edits, then train that behavior down with a two-term reward.

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
Hoang Cuong Nguyen, Mark Dras and Usman Naseem at Macquarie University compare supervised fine-tuning, reasoning-augmented fine-tuning and ORPO across Llama-3.1-8B, Gemma-2-9B and Qwen3-8B, and find that the choice of post-training method, not only the safety data, determines how refusal is computed inside the model.

When Models Edit Too Much: On the Fidelity of Minimal Code Edits
Tongyao Zhu, Wei Hern Lim and Min-Yen Kan at the National University of Singapore define over-editing as a measurable failure of code repair, build a controlled benchmark from 400 BigCodeBench problems, and show that edit fidelity is a separate axis from correctness that post-training can improve.

Rethinking On-Policy Distillation of Large Language Models II: One Training Example
Zixuan Fu and colleagues at Tsinghua push on-policy distillation to the data-minimal limit by training on a single query, and find it recovers most of full-data OPD's gain, which reframes what OPD is actually short of.

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
Boyan Li and colleagues show that the standard practice of fusing on-policy distillation and RLVR inside a single training step is worse than simply running them in sequence, and explain why with pass@k and coverage analysis.

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents
Yanting Yang, Can Jin, Dimitris Metaxas and colleagues (Rutgers) propose SPACE, which lets a long-horizon agent emit variable-length action chunks by distilling chunk boundaries from programmatic skills induced out of successful trajectories.

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
Jianlyu Chen, Hongjin Qian, Zheng Liu and a large BAAI-led team introduce DisCo and the AREX-Skill Library, distilling 1,000 widely used ML repositories into more than 5,000 verified reusable skills and showing that operational know-how, not the harness, is what research agents are missing.

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance
Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash and Akshay Krishnamurthy (Microsoft Research) ask whether standard SFT produces the model you actually want to run RL on, and propose a one-line change that says no.

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
Siye Wu and colleagues compare three ways to consolidate domain-expert RLVR models, Merge of task vectors, Mix RL of pooled datasets, and multi-teacher on-policy distillation, using shared experts and data across scales.

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs
Jian Wang and colleagues describe a production product-linking cascade at marketplace scale where a distilled cross-encoder auto-resolves the easy majority and an agentic VLM with web search settles only the ambiguous tail.

PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Hands an agent one base model, one GPU and ten hours, then asks it to post-train the model on its own. It measures AI automating AI training more directly than any other benchmark here, and audits the reward hacking that follows.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
Formalizes self-improvement around the generation-verification gap, the gain a model gets from filtering its own output with its own judgment. It gives the working rule for this collection, which is that a system improves itself only where it verifies better than it generates.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Evolves task prompts using mutation prompts, and evolves the mutation prompts too, so the instructions for improving prompts improve alongside them. It is the first system in the language-model era whose improvement operator is itself under optimization.