🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
392 papers · TrainingClear filters →
FrogNano: Training a 4B Coding Agent via Online Task Synthesis

FrogNano: Training a 4B Coding Agent via Online Task Synthesis

Small coding agents are usually built by distilling a frontier model's trajectories. Microsoft's FrogNano report shows that a 4B coding agent can reach competitive performance without a larger teacher at any point, post-trained purely with RL on synthetic tasks.

49Code
Miles v0.1: Production-Level Post-Training

Miles v0.1: Production-Level Post-Training

RadixArk releases Miles v0.1, an open-source post-training system built on the slime design, covering RL, LoRA RL, on-policy distillation and SFT, with an end-to-end case study running fully asynchronous agentic RL on a 744B-A40B GLM-5.2 model over terminal-use coding tasks.

50Training
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

The NeoHorse Team releases NeoHorse-1, a family of agent-native 4B and 9B models built on an agentic post-training loop in which a router's records of predicted capability demand and selected service tier become the training data for the next round.

51Agents
Optimizer Memory Schedules for Outscaling the Overtraining Axis

Optimizer Memory Schedules for Outscaling the Overtraining Axis

Katie Everett (MIT CSAIL) and Shikai Qiu (NYU) show that optimizer rankings and optimal hyperparameters change substantially as the training horizon extends, and argue that overtraining factor belongs on the axis list for optimizer evaluation.

52Training
What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

Zhinan Hou and colleagues at Tsinghua University and Meituan study what data on-policy distillation actually needs, and find that eight hard examples match a 17K-example baseline.

53Training
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Sanyuan Chen and colleagues at FAIR at Meta present Text-AB, a 3B latent-diffusion speech model that removes the forced-alignment stage from the Audiobox line and handles dubbing and two-speaker dialogue in one system.

54Safety
Improving precipitation forecasts in an AI weather model using observational data

Improving precipitation forecasts in an AI weather model using observational data

Julian F. Schmitt and colleagues at X, The Moonshot Factory and Google, with Caltech and Stanford, fine-tune an AI weather model on observed precipitation rather than reanalysis and improve precipitation forecasts substantially.

55Training
Normalized Low-Rank Adaptation

Normalized Low-Rank Adaptation

Jiale Kang and colleagues at CUHK, Yuanshi Intelligence and Microsoft Research normalize the down-projection matrices in LoRA and report faster convergence, better stability and less forgetting at no extra cost.

56Training
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Jacqueline He and colleagues at Meta AI, the University of Washington and Princeton show that standard knowledge distillation helps reasoning and hurts factual recall during mid-training, and trace the cause to teacher confidence.

57Training
Instruction Duplication as an Inference-Time Control Primitive

Instruction Duplication as an Inference-Time Control Primitive

Victor Lavrenko at PeaceTech VC measures instruction duplication, repeating only the procedural instruction, as a black-box inference-time control across seven instruction-tuned models and 16,800 scheduled generations, and reports gains on process compliance without any change in final-answer accuracy.

58Training
FailBench: How Reliable are VLMs at Judging Robot Task Success?

FailBench: How Reliable are VLMs at Judging Robot Task Success?

Zaruhi Navasardyan, Tatul Danielyan and Hrant Davtyan at Metric AI Lab assemble 2,197 real manipulation attempts from 14 sources and find that vision-language models used as robot success detectors reach only 0.77 mean balanced accuracy, with fine-tuned detectors doing worse than general-purpose models.

59Evaluation
CROCODIL: Cross-Model Code Editing with LLMs

CROCODIL: Cross-Model Code Editing with LLMs

Linghan Zhong, Aditya Thimmaiah, Milos Gligoric and Junyi Jessy Li at UT Austin with Cisco Research show that a model editing code originally written by a different model makes more and larger edits, then train that behavior down with a two-term reward.

60Training
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Hoang Cuong Nguyen, Mark Dras and Usman Naseem at Macquarie University compare supervised fine-tuning, reasoning-augmented fine-tuning and ORPO across Llama-3.1-8B, Gemma-2-9B and Qwen3-8B, and find that the choice of post-training method, not only the safety data, determines how refusal is computed inside the model.

61Training
When Models Edit Too Much: On the Fidelity of Minimal Code Edits

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

Tongyao Zhu, Wei Hern Lim and Min-Yen Kan at the National University of Singapore define over-editing as a measurable failure of code repair, build a controlled benchmark from 400 BigCodeBench problems, and show that edit fidelity is a separate axis from correctness that post-training can improve.

62Code
Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Zixuan Fu and colleagues at Tsinghua push on-policy distillation to the data-minimal limit by training on a single query, and find it recovers most of full-data OPD's gain, which reframes what OPD is actually short of.

63Training
Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

Boyan Li and colleagues show that the standard practice of fusing on-policy distillation and RLVR inside a single training step is worse than simply running them in sequence, and explain why with pass@k and coverage analysis.

64Training
Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

Yanting Yang, Can Jin, Dimitris Metaxas and colleagues (Rutgers) propose SPACE, which lets a long-horizon agent emit variable-length action chunks by distilling chunk boundaries from programmatic skills induced out of successful trajectories.

65Agents
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Jianlyu Chen, Hongjin Qian, Zheng Liu and a large BAAI-led team introduce DisCo and the AREX-Skill Library, distilling 1,000 widely used ML repositories into more than 5,000 verified reusable skills and showing that operational know-how, not the harness, is what research agents are missing.

66Training
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash and Akshay Krishnamurthy (Microsoft Research) ask whether standard SFT produces the model you actually want to run RL on, and propose a one-line change that says no.

67Training
Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

Siye Wu and colleagues compare three ways to consolidate domain-expert RLVR models, Merge of task vectors, Mix RL of pooled datasets, and multi-teacher on-policy distillation, using shared experts and data across scales.

68Training
Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

Jian Wang and colleagues describe a production product-linking cascade at marketplace scale where a distilled cross-encoder auto-resolves the easy majority and an agentic VLM with web search settles only the ambiguous tail.

69Agents
PostTrainBench: Can LLM Agents Automate LLM Post-Training?

PostTrainBench: Can LLM Agents Automate LLM Post-Training?

Hands an agent one base model, one GPU and ten hours, then asks it to post-train the model on its own. It measures AI automating AI training more directly than any other benchmark here, and audits the reward hacking that follows.

70Training
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Formalizes self-improvement around the generation-verification gap, the gain a model gets from filtering its own output with its own judgment. It gives the working rule for this collection, which is that a system improves itself only where it verifies better than it generates.

71Training
Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

Evolves task prompts using mutation prompts, and evolves the mutation prompts too, so the instructions for improving prompts improve alongside them. It is the first system in the language-model era whose improvement operator is itself under optimization.

72Training
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026