Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Fanchao Chen and Shivaram Venkataraman (UW-Madison) with Haibin Lin and colleagues at ByteDance Seed show why RL training with FP8 in both rollout and training collapses, and fix it with Calibrated Clipping.
Ask this paper
Instability traced. Full-pipeline FP8 RL shows mid-training entropy surges and garbled outputs even with train-inference mismatch corrections such as TIS. Even blockwise scaling leaves the PPO clip fraction nearly 4x the BF16 level.
Cause. Small FP8 errors in token probabilities are amplified 1.7x to 2.9x in the importance ratio. This pushes negative-advantage tokens outside the trust region and zeroes their gradients, so bad outputs are not penalized. About 70% of late-stage over-clipping happens at the lower bound, and nearly 90% of it comes from high-entropy responses.
Calibrated Clipping. Sets the FP8 clipping bounds by matching the lower-bound clipping quantile of the BF16 distribution and rebalancing the upper bound.
Results. Across GRPO and DAPO and models from 8B to 32B, entropy surges disappear and blockwise FP8 reaches scores comparable to BF16.
Speed. Tensorwise FP8 training reaches up to 1.5x BF16 throughput and blockwise 10% to 20%, on top of the roughly 30% rollout speedup reported for FP8 generation.
Abstract
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compounded FP8 quantization noise distorts the importance ratio, disproportionately pushing negative-advantage tokens outside the trust region and erroneously zeroing out their gradients. As a result, pathological outputs are not properly penalized and accumulate over the course of training. To address this, we propose Calibrated Clipping, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly. Extensive experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, and multiple FP8 scaling granularities demonstrate that our approach successfully eliminates entropy surges and restores performance comparable to the BF16 baseline.