🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 3, 2026
Agents

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

First page
GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
The curator’s take

Geyi Yang, Zhongxiang Dai and colleagues at CUHK-Shenzhen, Tianjin University, HIT Shenzhen and ECNU present GUI-HARVEST, an automatic harness optimizer that improves GUI agents with frozen backbones by grounding failure diagnosis in screenshots and repeated runs.

Ask this paper

Key points
01

Grounded diagnosis. Model outputs and executed actions are aligned with before-and-after screenshots, so each finding is tied to a specific interface transition.

02

Repeated runs as one unit. Multiple runs of the same task are compared to separate behavior differences that change the outcome from execution noise.

03

Bounded edits with predictions. Verified findings are grouped into recurring failure patterns and mapped to bounded source-code edits; each edit records its predicted behavioral effect before evaluation, and the prediction is checked alongside task score.

04

Results. On OSWorld-Verified, held-out gains hold across six backbones; Qwen3-VL-32B-Instruct gains 12.33 points. A frozen optimized harness improves GPT-5 by 13.87 points on WindowsAgentArena at 50 steps with no further optimization.

05

Against prior optimizers. With the same backbone and starting harness, it beats Self-Harness and Meta-Harness.

Abstract

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.

Every Monday
Get next week’s papers.
Subscribe on Substack