OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Zhilin Wang, Yi Dong and colleagues at NVIDIA introduce OSWorld-Pro, a computer-use benchmark that scores agents on each subgoal along the way instead of only on the final file or screen state.
Ask this paper
Subgoal-level scoring. Over 300 tasks broken into more than 2,800 sequentially dependent subgoals, grounded in over 67,000 human annotations of model trajectories. Each subgoal is tagged with the application it happens in (OS, VS Code, terminal, LibreOffice and others).
Harder than OSWorld. The best configuration, Claude Opus 4.8 at max effort, reaches 77.7% overall, while the top Opus model scores 83.4% on OSWorld. Open-weight models drop much further: the best reaches 55.1%, and MiniMax M3 scores 28.9% here against 75.2% on OSWorld.
Judge validated against humans. A GPT-5.6-Sol max-effort judge agrees with human labels 94.1% of the time, close to the 95.6% agreement between independent human reviewers. It matches humans on identifying the targeted subgoal (97.0 vs 98.4%) but lags on judging feasibility (61.9 vs 91.6%).
Failure modes outcome scoring hides. Process scoring surfaces subgoal-irrelevant actions and click-based mistakes as major failure sources, which calls for different fixes than keyboard-input errors.
Training recipe over size. Within one model family, recipe differences produce larger gaps than parameter count (Qwen 3.8 27B at 32.1% vs Qwen 3.6 27B at 10.2%).
Abstract
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agents fail in various tasks, obfuscating critical insight for subsequent improvement. For instance, agents that err during keyboard inputs would require a different mitigation strategy from those that fail to precisely provide click-based inputs on the graphical UI. We introduce OSWorld-Pro: a set of over 300 tasks containing over 2800 subgoals to enable the procedural evaluation of CUAs grounded in over 67,000 human annotations. We use robust human-aligned LLM-Judges to evaluate the fulfillment of OSWorld-Pro subgoals and thereby reveal the progress that models make throughout a series of sequentially dependent subgoals. Our findings reveal that OSWorld-Pro is challenging even for state-of-the-art LLMs, with top performers like Claude Opus 5 achieving only 75.7% vs. 83.4% on OSWorld. Furthermore, we identify critical process-focused failure modes of various models (e.g. subgoal-irrelevant actions and click-based mistakes) to provide insights to improve performance and efficiency of CUAs.