ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

Zhang, Ding and colleagues (Alibaba Token Hub and Amap) propose ArenaFlow, an RL framework for open-ended agent tasks that turns tournament rankings of trajectories into step-level and skill-level credit.
Ask this paper
Tournament reward. Relative ranking across a tournament gives trajectory-level rewards where reliable scalar rewards are not available.
Reflective evaluation. Each comparison also yields pivotal success steps, reusable strategy skills and attribution of which retrieved skills were used.
Step credit. Trajectory advantages are propagated to high-confidence pivotal steps according to how deep the trajectory survived in the tournament.
Skill memory. Skill utility is estimated from group-level usage attribution, and a global skill memory is updated, pruned and retrieved as a prior for later exploration.
Abstract
Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by replacing pointwise scoring with relative preferences. However, they still compress rich comparative feedback into a single trajectory-level reward, obscuring decisive intermediate steps and preventing successful behaviors from being consolidated into reusable skills. We propose ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning. ArenaFlow leverages tournament-based relative ranking to derive trajectory-level reward signals. Each comparison is further equipped with structured reflective evaluation, which reveals three types of supervision: pivotal success steps, reusable strategy skills, and usage attribution of retrieved skills. At the step level, ArenaFlow propagates trajectory-level advantages to high-confidence pivotal steps according to tournament survival depth, enabling more targeted optimization of local reasoning behaviors. At the skill level, ArenaFlow estimates skill utility from group-level usage attribution and maintains a global skill memory through utility-aware updating, pruning, and retrieval. The resulting high-utility skills further serve as policy priors for future exploration. Extensive experiments validate ArenaFlow's effectiveness on open-ended agent tasks.