🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 20, 2026
Training · Evaluation

What Does Privileged Information Add to On-Policy Self-Distillation?

First page
What Does Privileged Information Add to On-Policy Self-Distillation?
The curator’s take

XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang and Tat-Seng Chua build a benchmark that holds the problem fixed while varying what the teacher sees, and find the privileged information adds much less than distillation itself.

Ask this paper

Key points
01

AMPLE-Math. 5,319 mathematical problems with six reasoning views sharing the same answer, so each privileged view can be compared against matched reference-free distillation.

02

Most of the gain is not from privilege. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B's improvement under thinking-enabled evaluation, in domain and on external benchmarks.

03

Where a reference does help. Evidence is modest in Qwen and strongest for a polished solution; complete traces add two percentage points in SmolLM3-3B at step 50.

04

Rollout format dominates. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both model families while problems, references and evaluation stay fixed.

05

Significance. The authors read on-policy self-distillation as improving access to capabilities already shared between direct-response and thinking-enabled inference, so a privileged reference is worth what it adds to that cross-mode transfer.

Abstract

On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B's improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Evidence for an additional reference benefit is modest in Qwen, strongest for a polished solution, whereas complete traces add two percentage points in SmolLM3-3B at step 50. These benefits depend on the student being trained. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both families while the problems, references, and evaluation stay fixed. Teacher profiles and matched loss interventions in Qwen further show that changing token-level supervision can leave student behavior largely unchanged. Together, these findings suggest that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference. The value of a privileged reference is what it adds to this cross-mode transfer, not how much of the solution it reveals.

Every Monday
Get next week’s papers.
Subscribe on Substack