Beat the long tail: Distribution-Aware Speculative Decoding for RL Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shao, Zelei, Srivatsa, Vikranth, Srivastava, Sanjana, Wu, Qingyang, Ariyak, Alpay, Wu, Xiaoxia, Patel, Ameen, Wang, Jue, Liang, Percy, Dao, Tri, Zhang, Ce, Zhang, Yiying, Athiwaratkun, Ben, Xu, Chenfeng, Wang, Junxiong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911271418331136
author Shao, Zelei
Srivatsa, Vikranth
Srivastava, Sanjana
Wu, Qingyang
Ariyak, Alpay
Wu, Xiaoxia
Patel, Ameen
Wang, Jue
Liang, Percy
Dao, Tri
Zhang, Ce
Zhang, Yiying
Athiwaratkun, Ben
Xu, Chenfeng
Wang, Junxiong
author_facet Shao, Zelei
Srivatsa, Vikranth
Srivastava, Sanjana
Wu, Qingyang
Ariyak, Alpay
Wu, Xiaoxia
Patel, Ameen
Wang, Jue
Liang, Percy
Dao, Tri
Zhang, Ce
Zhang, Yiying
Athiwaratkun, Ben
Xu, Chenfeng
Wang, Junxiong
contents Reinforcement learning(RL) post-training has become essential for aligning large language models (LLMs), yet its efficiency is increasingly constrained by the rollout phase, where long trajectories are generated token by token. We identify a major bottleneck:the long-tail distribution of rollout lengths, where a small fraction of long generations dominates wall clock time and a complementary opportunity; the availability of historical rollouts that reveal stable prompt level patterns across training epochs. Motivated by these observations, we propose DAS, a Distribution Aware Speculative decoding framework that accelerates RL rollouts without altering model outputs. DAS integrates two key ideas: an adaptive, nonparametric drafter built from recent rollouts using an incrementally maintained suffix tree, and a length aware speculation policy that allocates more aggressive draft budgets to long trajectories that dominate makespan. This design exploits rollout history to sustain acceptance while balancing base and token level costs during decoding. Experiments on math and code reasoning tasks show that DAS reduces rollout time up to 50% while preserving identical training curves, demonstrating that distribution-aware speculative decoding can significantly accelerate RL post training without compromising learning quality.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13841
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
Shao, Zelei
Srivatsa, Vikranth
Srivastava, Sanjana
Wu, Qingyang
Ariyak, Alpay
Wu, Xiaoxia
Patel, Ameen
Wang, Jue
Liang, Percy
Dao, Tri
Zhang, Ce
Zhang, Yiying
Athiwaratkun, Ben
Xu, Chenfeng
Wang, Junxiong
Machine Learning
Reinforcement learning(RL) post-training has become essential for aligning large language models (LLMs), yet its efficiency is increasingly constrained by the rollout phase, where long trajectories are generated token by token. We identify a major bottleneck:the long-tail distribution of rollout lengths, where a small fraction of long generations dominates wall clock time and a complementary opportunity; the availability of historical rollouts that reveal stable prompt level patterns across training epochs. Motivated by these observations, we propose DAS, a Distribution Aware Speculative decoding framework that accelerates RL rollouts without altering model outputs. DAS integrates two key ideas: an adaptive, nonparametric drafter built from recent rollouts using an incrementally maintained suffix tree, and a length aware speculation policy that allocates more aggressive draft budgets to long trajectories that dominate makespan. This design exploits rollout history to sustain acceptance while balancing base and token level costs during decoding. Experiments on math and code reasoning tasks show that DAS reduces rollout time up to 50% while preserving identical training curves, demonstrating that distribution-aware speculative decoding can significantly accelerate RL post training without compromising learning quality.
title Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
topic Machine Learning
url https://arxiv.org/abs/2511.13841