🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 25, 2026
Agents · Code

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

First page
WatchPoint: Executable User Feedback for Real-World Agentic Web Development
The curator’s take

Guanqun Yang and Xueqing Liu (Stevens Institute of Technology) with Wei Yang (UT Dallas) build WatchPoint, a simulated user that writes and runs diagnostic scripts against a live web app to tell a coding agent why its last attempt failed.

Ask this paper

Key points
01

Developer-style feedback. Instead of screenshots, LLM-judge scores or natural-language corrections, WatchPoint clicks, inspects computed styles and runs commands against the running application and returns structured observations for the retry.

02

Setting. Web-Bench: 50 multi-file web projects with 1,000 sequentially dependent tasks checked by deterministic end-to-end tests; Claude Code with Sonnet 4.6 reaches only 13.4% Pass@1.

03

Recovery. On the categories where it is enabled, WatchPoint recovers 38 of 66 retried tasks (57.6%) against 45.1% from test errors alone; human testers in a user study recover 54.5%.

04

Modest aggregate gain. Because tasks are sequential, overall Pass@2 rises only from 26.6% to 28.1%, although one early recovery unblocked 8 later tasks in a single project.

05

Capability gap. The feedback helps weaker coding models and can add noise for a model that already self-corrects well from test errors (GLM-5 reaches 44.5% Pass@2 unaided on its own runs).

Abstract

When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge scoring, or natural-language corrections, but few interact with the live application the way a developer would. We introduce WatchPoint, a simulated-user system that mimics real developer behavior by generating and executing diagnostic scripts against the running application, producing structured observations that guide the coding model's retry. Unlike prior approaches that target single-file edits or evaluate using non-executable metrics, we operate on Web-Bench, a benchmark of 50 multi-file web projects comprising 1,000 sequentially dependent tasks, verified by deterministic end-to-end tests. WatchPoint recovers 57.6% of the tasks it diagnoses, and a controlled user study confirms the simulation's realism: human testers achieve a comparable recovery rate (54.5%), providing evidence that automated diagnostic scripts can substitute for interactive human testing on sequential web development tasks. We further identify a pattern of capability gaps that governs when simulated-user feedback is helpful and when it should be withheld.

Every Monday
Get next week’s papers.
Subscribe on Substack