🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

The Geometry of On-Policy Distillation

Free while signed in. Answers cite the passages they came from.

First page
The Geometry of On-Policy Distillation
The curator’s take

On-policy distillation (OPD) has become one of the most discussed post-training recipes of the year, but it has mostly been treated as a black box sitting somewhere between supervised fine-tuning and RL. This paper opens it up, characterizing how OPD changes a model's weights at the level of parameter geometry, and argues OPD is not a midpoint between SFT and RLVR but its own distinct kind of update.

Key points
01

It touches fewer weights: Compared with SFT, OPD updates affect far fewer parameters and largely avoid the dominant principal directions of weight space, which helps explain its sample efficiency.

02

Early subspace locking: OPD's cumulative updates rapidly collapse into a narrow, low-dimensional subspace early in training, rather than spreading across many directions as SFT does.

03

That subspace is functionally sufficient: Constraining training to the early-formed subspace preserves OPD performance but substantially degrades SFT, showing the small subspace genuinely carries the useful signal rather than being an artifact.

04

Why it matters: Knowing where in weight space OPD does its work turns a popular but poorly understood recipe into something with a mechanistic account. That makes the method easier to reason about, combine with other objectives, and improve deliberately instead of by trial and error.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack