How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Yukun Zhang, Kemu Xu and Yishen Chen at CUHK and the University of Edinburgh separate what a harness contributes by pairing real task-specific plans against shuffled policy text matched in word count, across two Retail experiments and an Airline pilot in tau^2-bench.
Ask this paper
Guidance content is worth 7.17 points. Across 265 matched cells, Fixed plans beat word-count-matched Sham text by 7.17 percentage points of oracle-verified success (90 percent task-clustered bootstrap interval 1.15 to 13.36), concentrated in higher-complexity tasks.
A read-only verifier rejects 61 percent of invalid episodes. It also withholds 17 percent of correct ones, at under one cent of extra cost per episode.
Which component wins depends on liability. At a low cost of erroneous acceptance the planning gain dominates; at a high cost the verifier's avoided false passes dominate.
The verifier alone captures most of the benefit. A standalone verifier recovers nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.
The Sham control is the methodological contribution. Matching word count isolates guidance content from the mere presence of extra text in the prompt, which most harness comparisons leave confounded.
Abstract
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.