Nvidia's NeMo-DCR cuts a 1T-parameter RL weight sync from 87.5 minutes to 150 seconds
The paper targets agentic reinforcement learning setups where training and rollout run on separate clusters, so every policy update must be copied across before the next batch. Copying a full 1T checkpoint between two AWS regions took 87.5 minutes in the authors' measurements, while only about 1% of BF16 weights actually change per step. NeMo-DCR sends just those deltas, using compressible XOR masks for most changes and overwrites for the rest, and streams them via object storage or a relay tree. The authors report refits of 30B to 1T models running 12 to 40 times faster than a full-checkpoint baseline even at 3% and 5% change rates. Receivers end up with exactly the same bits as a dense copy, and partial writes can be retried after failures. For teams running RL post-training, weight synchronization is often the hidden bottleneck once rollouts move to cheaper or remote capacity. Faster, exact refits make cross-region and multi-cloud RL more practical without giving up reproducibility.