NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

Songlin Jiang and Mario Di Francesco (Aalto University) with Zhiyu Li, Terry Kong and colleagues at NVIDIA present NeMo-DCR, a bit-exact delta-compressed weight synchronization (refit) for agentic RL, released in NVIDIA NeMo RL.
Ask this paper
Problem. Agentic RL runs rollout on separate serving clusters, so each policy update must be copied there. A full 1T-parameter checkpoint takes 87.5 minutes to transfer between two AWS regions.
Observation. Only 0.6% to 1.2% of BF16 weight elements change per GRPO step across six models, so a full transfer spends 98% to 99% of its bytes on unchanged values.
Method. Fixed affine index mappings place changes in the checkpoint's canonical coordinates (over 96% of MoE weight bytes), XOR masks and overwrites carry the changes so receivers end with the same bits as a dense refit, and a joint commit plus retries recover from mid-refit failures.
Results. Refits of 30B to 1T models are 12x to 40x faster than a transport-only full-checkpoint reference even at 3% and 5% change rates. A 1T refit at 3% takes 150 s instead of 87.5 min; 120B takes 22.6 s instead of 750 s.
Encoding. Mixed XOR and overwrite encoding cuts Qwen3 payload bytes by 38% to 40% compared with overwrites alone.
Abstract
Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40$\times$ faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.