🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 27, 2026
Agents · Reinforcement Learning

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

First page
Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning
The curator’s take

Xincheng Yao (Shanghai Jiao Tong, intern at Tencent AI Platform) and colleagues propose GRAFT, which merges GRPO rollouts into a trajectory graph to estimate step-level advantages that follow the textbook definition.

Ask this paper

Key points
01

Bias in GRPO. Group-normalized advantages are reliable per response but biased per step, since failed trajectories can contain useful steps.

02

Trajectory graph. Rollouts are grafted into one graph where shared states merge; node values come from Bellman iteration and each edge's credit is the value difference.

03

Graph GAE. Extends GAE to the graph to reduce the effect of state-value estimation error.

04

Results. Consistent gains over GRPO and recent agentic RL algorithms on multi-turn agent benchmarks.

Abstract

Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.

Every Monday
Get next week’s papers.
Subscribe on Substack