Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Yan Zhang and colleagues at CAS IIE and Xiaomi MiLM Plus propose MaP (Masked Trajectory Prediction), which trains GUI agents on several navigation tasks at once by masking parts of a trajectory and predicting them.
Ask this paper
Problem with mixing. Directly mixing step-wise decision, state-action alignment and long-horizon planning data gives only marginal gains because the objectives conflict and the data are heterogeneous.
One objective. Multi-turn interactions are treated as trajectories; masking arbitrary components and predicting them turns all three tasks into the same objective.
Role-aware adapters. Each token is routed to a specialized representation space according to its role, which reduces interference from heterogeneous data.
Results. Zero-shot on AndroidControl MaP beats direct mixing by 2.8, 5.4 and 2.1 points on the three abilities; after post-training on Qwen2.5-VL it adds 3.0 and 2.7 SR on AndroidControl-High and -Low, and 2.6 and 5.4 points in Pass@1 and Pass@4 on AndroidWorld.
Venue. Evaluated on AndroidWorld, AndroidControl, GUI-Odyssey, AITZ and Mind2Web. Accepted to EMNLP 2026.
Abstract
Graphical User Interface (GUI) Agents autonomously interact with software to fulfill user requests, where GUI navigation stands out as the most critical and challenging capability. Mastering this capability demands a complex synergy of step-wise decision-making, state-action alignment, and long-horizon planning. While directly mixing these corresponding navigation tasks seems intuitive to simultaneously acquire these skills, such a direct combination is severely bottlenecked by inconsistent optimization objectives and profound data heterogeneity. To overcome these barriers, we propose the MaP (stands for ``\textbf{M}asked Tr\textbf{a}jectory \textbf{P}rediction''), a unified framework that seamlessly harmonizes divergent GUI navigation tasks. By modeling multi-turn GUI interactions as a trajectory and defining training objectives through component masking and prediction, MaP shifts the optimization from task-specific marginal distributions to a consistent objective. Furthermore, to handle the data heterogeneity across multiple navigation tasks, we design a role-aware adapter learning module that dynamically routes each token to a specialized representation space. Extensive experiments on five representative GUI navigation benchmarks demonstrate that MaP effectively mitigates gradient conflicts and significantly outperforms the direct mixture training, establishing a robust paradigm for multi-task GUI navigation.