🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Agents · Reinforcement Learning

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

First page
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
The curator’s take

Dongwon Jung, Muhao Chen, Varun Chandrasekaran, Jaron Lanier and colleagues at UC Davis and Microsoft (with UW and Purdue) introduce ProVer, which gives step-level credit in agentic GRPO only to segments that a judge proposes and rollouts then confirm.

Ask this paper

Key points
01

Judge proposes, rollouts verify. An agentic judge compares successful and failed trajectories in a group and names the segment likely responsible for the difference. ProVer then estimates that segment's advantage from the change in terminal success rate between continuations sampled before and after it.

02

Credit only where verified. Positive estimates are added to the GRPO advantages of the tokens inside the segment, so the judge decides where to check but never sets the reward itself.

03

Results. Across ALFWorld, WebShop and SearchQA, ProVer has the best average at both scales, a relative improvement over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B.

04

Cost. Checking one proposed segment per group adds modest generation overhead and still helps when the judge is not a frontier model.

Abstract

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.

Every Monday
Get next week’s papers.
Subscribe on Substack