RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Wei, Zhao, Yuheng, An, Dakai, Wu, Tianyuan, Cao, Lunxi, Xiong, Shaopan, Huang, Ju, Wang, Weixun, Yang, Siran, Su, Wenbo, Wang, Jiamang, Qu, Lin, Zheng, Bo, Wang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
by: Wang, Weixun, et al.
Published: (2025)
by: Wang, Weixun, et al.
Published: (2025)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
Complementary Reinforcement Learning
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
by: Zhao, Long, et al.
Published: (2026)
by: Zhao, Long, et al.
Published: (2026)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
by: Feng, Weigang, et al.
Published: (2025)
by: Feng, Weigang, et al.
Published: (2025)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
by: Wang, Chong, et al.
Published: (2026)
by: Wang, Chong, et al.
Published: (2026)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
Generalize Synchronization Mechanism: Specification, Properties, Limits
by: Chien, Chih-Wei, et al.
Published: (2023)
by: Chien, Chih-Wei, et al.
Published: (2023)
A Fast Confirmation Rule (aka Fast Synchronous Finality) for the Ethereum Consensus Protocol
by: Asgaonkar, Aditya, et al.
Published: (2024)
by: Asgaonkar, Aditya, et al.
Published: (2024)
Hamster: A Fast Synchronous Byzantine Fault Tolerance Protocol
by: Fu, Ximing, et al.
Published: (2024)
by: Fu, Ximing, et al.
Published: (2024)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
by: Hu, Qinghao, et al.
Published: (2025)
by: Hu, Qinghao, et al.
Published: (2025)
FedNS: A Fast Sketching Newton-Type Algorithm for Federated Learning
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Laminar: A Scalable Asynchronous RL Post-Training Framework
by: Sheng, Guangming, et al.
Published: (2025)
by: Sheng, Guangming, et al.
Published: (2025)
P-TimeSync: A Precise Time Synchronization Simulation with Network Propagation Delays
by: Dai, Wei, et al.
Published: (2024)
by: Dai, Wei, et al.
Published: (2024)
Robust and Scalable Renaming with Subquadratic Bits
by: Bai, Sirui, et al.
Published: (2025)
by: Bai, Sirui, et al.
Published: (2025)
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
by: Song, Jingwei, et al.
Published: (2026)
by: Song, Jingwei, et al.
Published: (2026)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
by: Cheng, Rongxin, et al.
Published: (2025)
by: Cheng, Rongxin, et al.
Published: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
by: Xie, Shaokang, et al.
Published: (2025)
by: Xie, Shaokang, et al.
Published: (2025)
SparseRL-Sync: Lossless Weight Synchronization with ~100x Less Communication
by: Hu, Lucas, et al.
Published: (2026)
by: Hu, Lucas, et al.
Published: (2026)
Empowering Distributed Training with Sparsity-driven Data Synchronization
by: Wang, Zhuang, et al.
Published: (2023)
by: Wang, Zhuang, et al.
Published: (2023)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
by: Qin, Ruoyu, et al.
Published: (2025)
by: Qin, Ruoyu, et al.
Published: (2025)
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
by: Ruan, Chaoyi, et al.
Published: (2026)
by: Ruan, Chaoyi, et al.
Published: (2026)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
by: Li, Haley, et al.
Published: (2026)
by: Li, Haley, et al.
Published: (2026)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
by: He, Yongjun, et al.
Published: (2025)
by: He, Yongjun, et al.
Published: (2025)
Thunderdome: Timelock-Free Rationally-Secure Virtual Channels
by: Avarikioti, Zeta, et al.
Published: (2025)
by: Avarikioti, Zeta, et al.
Published: (2025)
Large-Scale Metric Computation in Online Controlled Experiment Platform
by: Xiong, Tao, et al.
Published: (2024)
by: Xiong, Tao, et al.
Published: (2024)
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
by: Du, Yuxuan, et al.
Published: (2026)
by: Du, Yuxuan, et al.
Published: (2026)
FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
by: Zhao, Bohan, et al.
Published: (2025)
by: Zhao, Bohan, et al.
Published: (2025)
Distributed Online Rollout for Multivehicle Routing in Unmapped Environments
by: Weber, Jamison W., et al.
Published: (2023)
by: Weber, Jamison W., et al.
Published: (2023)
SHARE: Optimizing Secure Hub Allocation and Routing Efficiency in Payment Channel Networks
by: Yang, Lingxiao, et al.
Published: (2025)
by: Yang, Lingxiao, et al.
Published: (2025)
Similar Items
-
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025) -
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026) -
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025) -
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025) -
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)