Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Zeyu Kang, Xinquan Chen, Xuhong Wang and colleagues at Shanghai AI Laboratory train a computer-use agent to finish benign tasks, work around hazards when a safe path exists, and refuse harmful goals, using one joint SFT-then-RL recipe called SCOPE.
Ask this paper
Conditional policy. The target behavior depends on risk: complete ordinary tasks, continue safely past environmental hazards, and refuse when the goal is harmful or no safe path remains.
SCOPE-Gen data. An automated pipeline synthesizes verifiable capability tasks and injects one of five hazard types while keeping the user instruction and evaluator unchanged, yielding 2,199 capability, 460 safe-continuation and 81 refusal trajectories (SATraj-OS).
Training. SFT on all three trajectory types, then online RL for task completion with a conjunctive reward that gives full credit only when the task is completed without triggering a hazard.
Results from Qwen3.5-9B. SCOPE-SFT reaches 49.72% on OSWorld and 66.30% attack avoidance on OS-BLIND; SCOPE-RL raises OSWorld to 54.17% while keeping 64.30% avoidance, the best capability-safety harmonic mean (58.80%) among evaluated agents.
Ablation. Refusal trajectories produce most of the attack-avoidance gain, while risk-handling trajectories keep more task utility at similar avoidance.
Abstract
Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete ordinary benign tasks, avoid environmental hazards and continue when a safe completion path remains, and refuse when the goal is harmful or no safe path exists. To learn this conditional policy, we develop Safety and Capability Optimization for Policy Execution (SCOPE), which jointly post-trains a CUA for task-execution capability and safety-aware decision making. To provide aligned training data for this joint objective, we further introduce SCOPE-Gen, an automated pipeline that synthesizes verifiable capability tasks and converts them into paired environment-risk variants while preserving their original goals. Using the resulting tasks, we construct SATraj-OS, a trajectory dataset comprising capability demonstrations, safe continuations, and explicit refusals. SCOPE first learns from all three trajectory types through supervised fine-tuning and then further improves task completion through online reinforcement learning. Starting from Qwen3.5-9B, SCOPE-RL achieves a 54.17% task success rate on OSWorld and a 64.30% attack-avoidance rate on OS-BLIND, yielding the best aggregate capability--safety score of 58.80% among the evaluated agents. Ablations reveal asymmetric but complementary roles for the two forms of safety supervision: refusal trajectories account for most of the attack-avoidance gain, whereas risk-handling trajectories preserve greater task utility at comparable attack-avoidance levels.