RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gao, Wei, Zhao, Yuheng, An, Dakai, Wu, Tianyuan, Cao, Lunxi, Xiong, Shaopan, Huang, Ju, Wang, Weixun, Yang, Siran, Su, Wenbo, Wang, Jiamang, Qu, Lin, Zheng, Bo, Wang, Wei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909805907542016
author Gao, Wei
Zhao, Yuheng
An, Dakai
Wu, Tianyuan
Cao, Lunxi
Xiong, Shaopan
Huang, Ju
Wang, Weixun
Yang, Siran
Su, Wenbo
Wang, Jiamang
Qu, Lin
Zheng, Bo
Wang, Wei
author_facet Gao, Wei
Zhao, Yuheng
An, Dakai
Wu, Tianyuan
Cao, Lunxi
Xiong, Shaopan
Huang, Ju
Wang, Weixun
Yang, Siran
Su, Wenbo
Wang, Jiamang
Qu, Lin
Zheng, Bo
Wang, Wei
contents Reinforcement Learning (RL) is a pivotal post-training technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, synchronous RL post-training often suffers from significant GPU underutilization, referred to as bubbles, caused by imbalanced response lengths within rollout steps. Many RL systems attempt to alleviate this problem by relaxing synchronization, but this can compromise training accuracy. In this paper, we introduce tail batching, a novel rollout scheduling strategy for synchronous RL that systematically consolidates prompts leading to long-tail responses into a small subset of rollout steps (long rounds), while ensuring that the majority of steps (short rounds) involve only balanced, short rollouts. By excluding long responses from short rounds and rescheduling them into a few designated long rounds, tail batching effectively reduces GPU idle time during rollouts and significantly accelerates RL training without sacrificing accuracy. We present RollPacker, a system that fully harnesses the benefits of tail batching through holistic optimizations across all three RL stages: elastic parallelism adaptation for rollout, dynamic resource allocation and scheduling for reward, and stream-based training. Empirical results show that RollPacker achieves a 2.03x-2.56x end-to-end training time reduction compared to veRL and up to 2.24x speedup compared to RLHFuse for the Qwen2.5 family of LLMs on up to 128 H800 GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
Gao, Wei
Zhao, Yuheng
An, Dakai
Wu, Tianyuan
Cao, Lunxi
Xiong, Shaopan
Huang, Ju
Wang, Weixun
Yang, Siran
Su, Wenbo
Wang, Jiamang
Qu, Lin
Zheng, Bo
Wang, Wei
Distributed, Parallel, and Cluster Computing
Machine Learning
Reinforcement Learning (RL) is a pivotal post-training technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, synchronous RL post-training often suffers from significant GPU underutilization, referred to as bubbles, caused by imbalanced response lengths within rollout steps. Many RL systems attempt to alleviate this problem by relaxing synchronization, but this can compromise training accuracy. In this paper, we introduce tail batching, a novel rollout scheduling strategy for synchronous RL that systematically consolidates prompts leading to long-tail responses into a small subset of rollout steps (long rounds), while ensuring that the majority of steps (short rounds) involve only balanced, short rollouts. By excluding long responses from short rounds and rescheduling them into a few designated long rounds, tail batching effectively reduces GPU idle time during rollouts and significantly accelerates RL training without sacrificing accuracy. We present RollPacker, a system that fully harnesses the benefits of tail batching through holistic optimizations across all three RL stages: elastic parallelism adaptation for rollout, dynamic resource allocation and scheduling for reward, and stream-based training. Empirical results show that RollPacker achieves a 2.03x-2.56x end-to-end training time reduction compared to veRL and up to 2.24x speedup compared to RLHFuse for the Qwen2.5 family of LLMs on up to 128 H800 GPUs.
title RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2509.21009