🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Efficiency · Architecture · Reinforcement Learning

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

First page
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
The curator’s take

Xin Wang, Dayiheng Liu, Jianwei Zhang and colleagues at Alibaba Group (Qwen team, with Ohio State) present TRACE, an FP4 quantization framework for RL training of MoE language models that aligns training-side and rollout-side quantization.

Ask this paper

Key points
01

Problem. FP4 rollout cuts RL cost, but existing methods tune training-path and rollout-path quantization separately, leaving a mismatch between the two quantized executions that destabilizes RL.

02

Observation. For over 99% of mismatched quantized values, training and rollout differ by one adjacent FP4 codebook entry.

03

Method. Rollout-guided quantization-aware training uses rollout-side quantization results to set training-side FP4 rounding. A cache keeps mantissa and scale information only for deeper layers to limit storage and communication.

04

Results. On Qwen3.5-35B-A3B, Qwen3.5-122B-A10B, Qwen3.8-Flash-Next and Qwen3.8-2.4T-A95B across reasoning, coding and long-horizon tasks, joint FP4 weight, activation and KV-cache rollout matches BF16 rollout performance with up to 5.4x rollout speedup.

05

Overhead. Step time rises from 664 to 713 seconds, a 7.4% overhead over plain joint FP4 rollout; guidance accounts for 6% of rollout time.

Abstract

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.

Every Monday
Get next week’s papers.
Subscribe on Substack