🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 2, 2026
Reinforcement Learning

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

First page
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
The curator’s take

Chenyu Zhou and colleagues make an uncomfortable argument for multi-turn agentic RL: given a terminal-state verifier, spreading reward uniformly beats every attempt to target the turns that mattered, and they name the quantity that predicts when this holds.

Ask this paper

Key points
01

Verifier information density V_d = k/C: The fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes. Terminal-state verifiers sit deep in the low V_d regime where targeting is the wrong axis.

02

Targeting is second-order: In shared-rollout comparisons on tau^2-bench, continuous dense reward spread uniformly beats sparse binary outcome reward, while concentrating the same advantage on progress turns or on random turns is equally harmful.

03

The mechanism is coverage: Terminal-state verification collapses observable signal to a single final-write turn (k=1 in 98 percent of rollouts) while success requires a 5 to 8 step chain of prerequisite tool calls.

04

A phase boundary and a control everyone should run: Synthetic crossover at V_d about 0.8, against measured 0.15 on tau^2-bench and 0.4 on BFCL V3. The authors contribute a matched-concentration shuffled control that any targeting claim should have to clear, with 32 pre-registered seeds and an independent 20-seed replication.

05

Why it matters: This is a well-powered negative result aimed squarely at a popular research direction, and the pre-registration and shuffled control make it hard to wave away.

Abstract

Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes), and show that terminal-state verifiers sit deep in a low-V_d regime where targeting is the wrong axis. In controlled shared-rollout comparisons on tau^2-bench that separate reward density from credit geometry, a continuous dense reward spread uniformly beats the sparse binary outcome reward (net-harmful on 4/5 seeds), while concentrating the same advantage on progress turns or on random turns is equally harmful: targeting is second-order. The mechanism is coverage: terminal-state verification collapses the observable signal to a single final-write turn (k=1 in 98% of rollouts) while success requires a 5-8 step chain of prerequisite tool calls. A synthetic phase boundary places the crossover at V_d* ~ 0.8, whereas measured V_d is ~0.15 on tau^2-bench and ~0.4 on BFCL V3; uniform also wins on BFCL, where a matched-concentration shuffled control is negative on 8/8 seeds. The effect reproduces across model families on ToolACE-2-8B (Delta = -0.048 over 32 pre-registered seeds; an independent 20-seed replication is itself significant), and a pre-registered matched-budget breadth sweep traces a monotone dose-response whose deficit vanishes only at full chain coverage, with a reward-to-go arm reaching full-coverage parity. Uniform redistribution is the zero-information coverage default that per-turn schemes must beat; we contribute the matched-concentration shuffled control that any targeting claim should clear.

Every Monday
Get next week’s papers.
Subscribe on Substack