🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Evaluation · Agents

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

First page
WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness
The curator’s take

Yun-Yun Tsai (Columbia, during an internship at Meta) with Yuning Mao and colleagues at Meta Superintelligence Labs introduce WebUIProof, a benchmark that scores generated web interfaces by having a UI agent execute interaction tests in a headless browser.

Ask this paper

Key points
01

Benchmark. 219 tasks with structured specifications and element-aligned executable tests, covering general WebUIs such as dashboards, games and IDEs, and 3D simulation interfaces with physics-driven scenes. Specs and tests are built with an agent-assisted pipeline.

02

Evaluation protocol. A UI agent plans interactions, clicks, types and scrolls, observes DOM and visual changes, and checks assertions, so outputs that build and render but do not work under interaction are counted as failures.

03

Findings. Across eight commercial LLMs, the best on general WebUI is Claude Sonnet 4 at 46.10% accuracy, and it still fails to build in 30.37% of cases. 3D simulation tasks are harder for every model.

04

VisRL. The same harness provides RL rewards from build success, visual quality and test passes; training Qwen2.5-14B and MiMo-7B with it improves accuracy over SFT by 15.7% on general WebUI and 4.9% on 3D tasks.

Abstract

Evaluating WebUI code generation at scale is difficult: outputs may compile and look plausible yet fail under user interaction, and prior benchmarks largely rely on free-form prompts with static checks (build success, screenshots) that miss functional correctness. We introduce WebUIProof, an execution-oriented benchmark that provides structured specifications and dense, executable interaction tests for WebUI generation across two task families: general WebUIs (e.g., dashboards, game, interactive tools) and 3D interactive simulation (e.g., particle/galaxy systems, physics dynamics). WebUIProof includes a UI-agent harness that runs executable interaction tests in a headless browser using an iterative plan--act--observe loop: it locates DOM elements, performs actions, observes resulting UI/DOM changes, and checks the specified assertions. We evaluate across eight commercial LLMs and observe frequent failures on interaction-based requirements even when pages render successfully, especially on 3D simulation interfaces. Finally, we show the UI-agent harness can provide outcome-level training signals. Training compact models (e.g., Qwen2.5 14B and MIMO 7B) with RL rewards derived from executable interaction tests improves functional completion while reducing build failures.

Every Monday
Get next week’s papers.
Subscribe on Substack