Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions

Xunjian Yin, Bhuwan Dhingra, Shuyan Zhou and colleagues at Duke with Amazon collaborators build BreakingWeb, which makes browser tasks harder by changing the environment under a task agents already solve while keeping the instruction and success criterion fixed.
Ask this paper
Construction. 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families. Each intervention changes one web-stack layer, is deterministic, detectable and recoverable, and is graded on the final backend state.
Agents lose much more than humans. Interventions cut agent pass rate by 22.9% on average and overturn nearly half the tasks each agent solved cleanly. Humans lose 10.0% on a first attempt and 5.7% after one familiarization attempt.
Belief failure dominates. 75% of failures across the six browser-use agents end with the agent declaring success although the required state change never happened.
Two harnesses, two failure modes. Text-based browser-use agents need to verify external state before trusting their own done signal, while screenshot-only GUI agents mostly stall because they cannot find the affordance on the rendered page.
Models tested. Opus-4.7, Sonnet-4.6, GPT-5.4, GPT-5.4-mini, Gemini-3.1-Pro and Gemini-3-Flash, plus GUI-only variants and two open-weight agents.
Abstract
As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challenging. We instead construct challenging instances from tasks agents already solve, turning difficulty into a programmable property of the environment. BreakingWeb pairs every base task with an intervention condition that preserves the user instruction, latent target, and backend success criterion while changing the environment at different web stack layers. Each intervention is deterministic, detectable, and recoverable, and is annotated with the cognitive primitive it primarily loads. The benchmark contains 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families, all graded against outcomes. We evaluate six strong browser-use agents, three GUI-only agents that see only screenshots, and humans. The construction is effective: interventions cut agent pass rate by 22.9% on average and overturn nearly half of the tasks each agent solves cleanly, whereas humans lose 10.0% on a first attempt and 5.7% after one familiarisation attempt. The dominant failure is belief failure: 75% of the six agents' failures end with a declared success although the required change never happened. Our code, data and environment are publicly available at www.breakingweb.app.