Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

Yan Zhang, Yu Zhou and colleagues at the Institute of Information Engineering (CAS), UCAS, Tencent and Tsinghua present GUI-SD-v2, which extends on-policy self-distillation from GUI grounding to multi-turn GUI interaction.
Ask this paper
Starting point. On-policy self-distillation gives dense token-level supervision from a copy of the model that sees privileged information, and it worked for single-step grounding.
Two obstacles. In multi-turn settings the self-teacher follows privileged guidance poorly, and the guidance itself is too thin.
Stage one. Jointly optimize rollouts with and without privileged guidance from the same GUI states so the model learns to follow that guidance.
Stage two. Selectively distill step-specific reasoning and memory guidance through the privilege-conditioned self-teacher, so the student learns what to do and what to remember for later steps.
Results. On AndroidWorld and MobileWorld it compares favorably with self-distillation baselines and beats the evaluated state-of-the-art methods on Pass@1 and Pass@3.
Abstract
Graphical User Interface (GUI) agents enable the fulfillment of complex user instructions through multi-turn interactions with software environments, requiring step-wise reasoning and long-horizon memory to guide actions and retain task-relevant information, respectively. Recent on-policy self-distillation (OPSD) methods have achieved strong performance on GUI grounding, a foundational subtask for GUI agents, owing to dense token-level supervision from privilege-conditioned self-teachers. However, extending existing OPSD methods to multi-turn GUI agents is hindered by self-teachers' limited privilege-following ability and insufficient privileged guidance. In this paper, we introduce GUI-SD-v2, the next version of GUI-SD, which extends OPSD from GUI grounding to multi-turn GUI interaction and addresses key limitations through a two-stage training framework. Specifically, GUI-SD-v2 first strengthens privilege following by jointly optimizing rollouts with and without privileged guidance from the same GUI states. Furthermore, it selectively distills step-specific reasoning and memory guidance through a privilege-conditioned self-teacher, supporting action decisions and the retention of task-relevant information for subsequent interactions. Extensive experiments on two representative GUI agent benchmarks, AndroidWorld and MobileWorld, show that GUI-SD-v2 compares favorably with existing OPSD baselines while consistently outperforming the evaluated state-of-the-art methods in both Pass@1 and Pass@3 success rates. Code and training data will be publicly released.