🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Agents · Reinforcement Learning

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

First page
PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
The curator’s take

Jiaan Zhu, Wei Gao and colleagues at USTC and HKUST with Alibaba Group present PEARL, an asynchronous agentic RL system that combines elastic GPUs, temporary reuse of idle training GPUs, and per-workload choice between prefill-decode colocation and disaggregation to speed up multi-turn rollouts.

Ask this paper

Key points
01

Problem. In a Qwen3-8B SWE-bench run, rollout still takes 81.8% of end-to-end agentic RL time with four times the training GPUs. Going from 16 to 24 rollout GPUs cuts rollout time by only 10.3%.

02

Prefill-decode choice depends on load. Disaggregating prefill and decode beats colocation by 9.7% at batch size 32, but the advantage falls to 1.4% at 80 and colocation wins by 6.1% at 96, so the configuration has to adapt.

03

System. PEARL tracks GPUs, workers and the prefill-decode layout together, picks configurations that shorten batch completion under the current GPU budget, and lends idle training GPUs to rollout between updates, gated by an estimate of switching cost.

04

Results. Against RLBoost+ on the same GPU availability traces, PEARL raises throughput by up to 26.9% for Qwen3-8B and 36.3% for Qwen3-30B-A3B, and by 5.3% to 63.7% across tested batch sizes.

Abstract

Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$--$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.

Every Monday
Get next week’s papers.
Subscribe on Substack