Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914516982300672 |
|---|---|
| author | Iso, Hayate Mitra, Tiyasa Mondal, Sudipta Shafipour, Rasoul Elango, Venmugil Kong, Terry Huang, Yuki Na, Seonjin Putterman, Izzy Chislett, Benjamin Ashkenazi, Maor Guman, Joseph Shen, Gerald Konuk, Tugrul Aithal, Ashwath Borkar, Ritika Zilberstein, Ran Rouhani, Bita |
| author_facet | Iso, Hayate Mitra, Tiyasa Mondal, Sudipta Shafipour, Rasoul Elango, Venmugil Kong, Terry Huang, Yuki Na, Seonjin Putterman, Izzy Chislett, Benjamin Ashkenazi, Maor Guman, Joseph Shen, Gerald Konuk, Tugrul Aithal, Ashwath Borkar, Ritika Zilberstein, Ran Rouhani, Bita |
| contents | RL post-training of frontier language models is increasingly bottlenecked by autoregressive rollout generation, making rollout acceleration a central systems challenge. Many existing efficiency methods improve throughput by changing the rollout or optimization regime, for example, through off-policy execution, replay, or lower-precision generation. We study speculative decoding as a lossless acceleration primitive for RL rollouts that preserves the target model's output distribution. We implement speculative decoding in NeMo-RL with a vLLM backend, supporting both synchronous and asynchronous pipelines and enabling speculation during RL rollouts. This benefit is realizable across speculation mechanisms, such as pretrained MTP heads, small external draft models or even techniques such as Eagle3, which are traditionally applied after RL phase. This yields a deployment path for state-of-the-art speculative decoding inside RL training. In a reasoning post-training workload at 8B scale under synchronous RL, speculative decoding improves rollout throughput by 1.8x. Using a high-fidelity performance simulator, we project that combining speculative decoding with asynchronous RL yields up to 2.5x end-to-end training speedup at 235B scale. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_26779 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding Iso, Hayate Mitra, Tiyasa Mondal, Sudipta Shafipour, Rasoul Elango, Venmugil Kong, Terry Huang, Yuki Na, Seonjin Putterman, Izzy Chislett, Benjamin Ashkenazi, Maor Guman, Joseph Shen, Gerald Konuk, Tugrul Aithal, Ashwath Borkar, Ritika Zilberstein, Ran Rouhani, Bita Machine Learning Computation and Language RL post-training of frontier language models is increasingly bottlenecked by autoregressive rollout generation, making rollout acceleration a central systems challenge. Many existing efficiency methods improve throughput by changing the rollout or optimization regime, for example, through off-policy execution, replay, or lower-precision generation. We study speculative decoding as a lossless acceleration primitive for RL rollouts that preserves the target model's output distribution. We implement speculative decoding in NeMo-RL with a vLLM backend, supporting both synchronous and asynchronous pipelines and enabling speculation during RL rollouts. This benefit is realizable across speculation mechanisms, such as pretrained MTP heads, small external draft models or even techniques such as Eagle3, which are traditionally applied after RL phase. This yields a deployment path for state-of-the-art speculative decoding inside RL training. In a reasoning post-training workload at 8B scale under synchronous RL, speculative decoding improves rollout throughput by 1.8x. Using a high-fidelity performance simulator, we project that combining speculative decoding with asynchronous RL yields up to 2.5x end-to-end training speedup at 235B scale. |
| title | Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2604.26779 |