Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents

Tong Zhang, Yihao Liu, Jiahua Bao and colleagues at Alibaba's Qwen Large Model Application Team, with Peking University and other universities, show that token-level teacher-student gaps are a poor guide to where on-policy distillation helps multi-turn agents, and propose OG-OPD, which reweights supervision using the student's actual task outcomes.
Ask this paper
Supervision-benefit mismatch. On ScienceWorld and WebShop, swapping in the teacher's action at a large-gap turn makes the outcome worse (teacher harm) in 10.8% of decisions and better (teacher rescue) in only 8.8%. Among small-gap turns, 7.6% still change the final outcome.
Method. OG-OPD flags candidate turns from the current gap and sudden gap increases, then runs paired continuations of the student from its own action and from the teacher's action. Supervision is increased only where the teacher action improves the student's final result.
Results. With a Qwen3-32B teacher and Qwen3-1.7B student, OG-OPD reaches 45.7% success on WebShop against 28.0% for vanilla OPD and 38.7% for the strongest baseline, and 64.2% on ALFWorld against 62.7%.
Selective by design. Candidate turns are about a quarter of eligible turns, but fewer than 1% of recorded turns receive extra weight. Upweighting candidate turns without the outcome check reaches only 35.0% on WebShop.
Abstract
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.