Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Iso, Hayate, Mitra, Tiyasa, Mondal, Sudipta, Shafipour, Rasoul, Elango, Venmugil, Kong, Terry, Huang, Yuki, Na, Seonjin, Putterman, Izzy, Chislett, Benjamin, Ashkenazi, Maor, Guman, Joseph, Shen, Gerald, Konuk, Tugrul, Aithal, Ashwath, Borkar, Ritika, Zilberstein, Ran, Rouhani, Bita
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914516982300672
author Iso, Hayate
Mitra, Tiyasa
Mondal, Sudipta
Shafipour, Rasoul
Elango, Venmugil
Kong, Terry
Huang, Yuki
Na, Seonjin
Putterman, Izzy
Chislett, Benjamin
Ashkenazi, Maor
Guman, Joseph
Shen, Gerald
Konuk, Tugrul
Aithal, Ashwath
Borkar, Ritika
Zilberstein, Ran
Rouhani, Bita
author_facet Iso, Hayate
Mitra, Tiyasa
Mondal, Sudipta
Shafipour, Rasoul
Elango, Venmugil
Kong, Terry
Huang, Yuki
Na, Seonjin
Putterman, Izzy
Chislett, Benjamin
Ashkenazi, Maor
Guman, Joseph
Shen, Gerald
Konuk, Tugrul
Aithal, Ashwath
Borkar, Ritika
Zilberstein, Ran
Rouhani, Bita
contents RL post-training of frontier language models is increasingly bottlenecked by autoregressive rollout generation, making rollout acceleration a central systems challenge. Many existing efficiency methods improve throughput by changing the rollout or optimization regime, for example, through off-policy execution, replay, or lower-precision generation. We study speculative decoding as a lossless acceleration primitive for RL rollouts that preserves the target model's output distribution. We implement speculative decoding in NeMo-RL with a vLLM backend, supporting both synchronous and asynchronous pipelines and enabling speculation during RL rollouts. This benefit is realizable across speculation mechanisms, such as pretrained MTP heads, small external draft models or even techniques such as Eagle3, which are traditionally applied after RL phase. This yields a deployment path for state-of-the-art speculative decoding inside RL training. In a reasoning post-training workload at 8B scale under synchronous RL, speculative decoding improves rollout throughput by 1.8x. Using a high-fidelity performance simulator, we project that combining speculative decoding with asynchronous RL yields up to 2.5x end-to-end training speedup at 235B scale.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26779
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
Iso, Hayate
Mitra, Tiyasa
Mondal, Sudipta
Shafipour, Rasoul
Elango, Venmugil
Kong, Terry
Huang, Yuki
Na, Seonjin
Putterman, Izzy
Chislett, Benjamin
Ashkenazi, Maor
Guman, Joseph
Shen, Gerald
Konuk, Tugrul
Aithal, Ashwath
Borkar, Ritika
Zilberstein, Ran
Rouhani, Bita
Machine Learning
Computation and Language
RL post-training of frontier language models is increasingly bottlenecked by autoregressive rollout generation, making rollout acceleration a central systems challenge. Many existing efficiency methods improve throughput by changing the rollout or optimization regime, for example, through off-policy execution, replay, or lower-precision generation. We study speculative decoding as a lossless acceleration primitive for RL rollouts that preserves the target model's output distribution. We implement speculative decoding in NeMo-RL with a vLLM backend, supporting both synchronous and asynchronous pipelines and enabling speculation during RL rollouts. This benefit is realizable across speculation mechanisms, such as pretrained MTP heads, small external draft models or even techniques such as Eagle3, which are traditionally applied after RL phase. This yields a deployment path for state-of-the-art speculative decoding inside RL training. In a reasoning post-training workload at 8B scale under synchronous RL, speculative decoding improves rollout throughput by 1.8x. Using a high-fidelity performance simulator, we project that combining speculative decoding with asynchronous RL yields up to 2.5x end-to-end training speedup at 235B scale.
title Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2604.26779