🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 7, 2026
Agents · Reasoning

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

First page
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
The curator’s take

Yifan Zhang, Yutong Dai, Ran Xu, Zeyuan Chen and colleagues at Salesforce AI Research introduce CLIFT, which trains a web agent to verify its own rollouts and reuses that verifier for test-time trajectory selection without an external judge.

Ask this paper

Key points
01

Training signal. The agent answers natural-language verification questions about its rollouts. A conformal certifier keeps only questions whose evidence agrees with a training-time judge, weights them by signed trust, and adds the score to per-step rewards without ever subtracting from the judge baseline.

02

WebArena Infinity. A trained Gemma-4 (31B) agent with Conformal Trajectory Selection reaches 74.6% over 9 apps, 12.8 points above the same base model and above Gemini 3 Flash with browser-use at 70.1%, winning 7 of 9 apps.

03

Transfer to a closed model. On VisualWebArena, the certified question bank from the open model is applied to GPT-5.5 at test time and reaches state of the art under the canonical harness.

04

Zero-shot. On Online Mind2Web, no agent is trained; a translated question bank alone improves a live-web agent.

Abstract

Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.

Every Monday
Get next week’s papers.
Subscribe on Substack