🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 3, 2026
Agents

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

First page
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
The curator’s take

Caiqi Zhang, Rujun Han, Chen-Yu Lee, Tomas Pfister and colleagues at Google Cloud AI Research and the University of Cambridge introduce VeriHarness, which turns the generator's own base model into an agentic verifier for long-horizon workspace tasks without reference answers or rubrics.

Ask this paper

Key points
01

Observation. Across repeated rollouts, disagreement often exposes correct alternatives, while consensus can hide shared errors.

02

Two verifier roles. A disagreement resolver checks competing claims against evidence in the workspace, and a consensus challenger tests claims the rollouts agree on and looks for omitted requirements.

03

Results. Across five long-horizon benchmarks VeriHarness gives the best selection scores among baselines; with evidence-backed revision the gain over a single rollout is 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.

04

Release. Verification skills self-improve from failure feedback, and the authors release about 26,000 rollouts that cost over $100,000 to produce.

Abstract

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.

Every Monday
Get next week’s papers.
Subscribe on Substack