ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Seil Kang, Hangoo Kang, Tarun Suresh and Azalia Mirhoseini (Stanford) with Yonsei, Korea University and Bespoke Labs present ThunderSyncRL, which streams gradient computation during agentic RL without introducing policy staleness.
Ask this paper
Idea. Gradient work starts as soon as its inputs are fixed. For GRPO, each trajectory's score gradient is computed when its reward arrives instead of waiting for the full group. For on-policy distillation, gradients for a completed turn are computed while the next tool calls run in the sandbox.
Exactness. The authors prove that gradient streaming yields the same GRPO and OPD updates as batch-synchronous training, so neither objective changes.
Speed. On SWE-bench Verified and Terminal Bench 4.0, models reach the same performance up to 1.9x faster than synchronous training.
Against async training. Asynchronous pipelines remove idle time by accepting stale policies; at a fixed budget, ThunderSyncRL beats asynchronous training by up to 2.47 percentage points.
Abstract
Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to $1.9 \times$ faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to $2.47$ percentage points.