🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 26, 2026
Agents · Safety

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

First page
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
The curator’s take

Yuxuan Li (Carnegie Mellon, interning at Microsoft Research) with Will Epperson, Wesley Deng and Zezhou Huang of Microsoft Research build CAVEAT, a benchmark that tests whether computer-use agents still buy the product that is best for the user when the marketplace has its own incentives.

Ask this paper

Key points
01

Setup. Nine browser marketplaces (lodging, retail, food delivery, resale, freelance, grocery, specialty) and eight steering mechanisms taken from documented commercial practice, including sponsored placement, preferential ranking and drip pricing. The user request and catalog stay fixed while the mechanisms are switched on or off.

02

Main result. Across five model families on 52 Standard tasks, optimal purchases fall from 78.6% in the matched control to 17.3% with steering enabled. On the 2,000-plus-product Hard split, GPT-5.6-Sol at high reasoning goes from 90.0% to 0.0%, and all 50 of its failures stop on page 1 of an 88-page catalog.

03

Single mechanism is enough. With only drip pricing active, GPT-5.6-Terra makes 35 suboptimal purchases in 60 runs, every one of them the drip-priced promoted product.

04

Diagnosis. Steering enters at three points. Agents reweight the user's priorities, narrow the candidate set too early, and commit before checking decision-relevant evidence.

05

Fixes. CAVEAT-Harness (structured objective plus a verification tool) lifts GPT-5.6-Terra from 11.7% to 66.7% and GPT-5.6-Sol-high on Hard from 0.0% to 80.0%; the verification tool alone reaches 66.7%. Qwen3.5-27B gains little from the harness (4.2%), but post-training it on harness trajectories (CAVEAT-27B) reaches 22.9%.

Abstract

Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives that do not align with the user's? In online marketplaces, for example, platforms may favor some products over others, potentially steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and increased reasoning improve robustness, but substantial failures persist. Our trajectory analysis and targeted ablations identify three points where steering enters the decision process: (1) agents distort the user's priorities, (2) prematurely narrow the set of alternatives they consider, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which directly targets these failure modes and raises user-optimal purchasing by 55.0%. Targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose how it fails, and show that targeted interventions can substantially improve it.

Every Monday
Get next week’s papers.
Subscribe on Substack