Zero Bubble Pipeline Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Penghui, Wan, Xinyi, Huang, Guangxing, Lin, Min |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
by: Wan, Xinyi, et al.
Published: (2025)
by: Wan, Xinyi, et al.
Published: (2025)
Pipeline Parallelism with Controllable Memory
by: Qi, Penghui, et al.
Published: (2024)
by: Qi, Penghui, et al.
Published: (2024)
Revisiting Parameter Server in LLM Post-Training
by: Wan, Xinyi, et al.
Published: (2026)
by: Wan, Xinyi, et al.
Published: (2026)
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024)
by: Yeung, Man Tsung, et al.
Published: (2024)
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
by: Guo, Jihu, et al.
Published: (2025)
by: Guo, Jihu, et al.
Published: (2025)
FreeRide: Harvesting Bubbles in Pipeline Parallelism
by: Zhang, Jiashu, et al.
Published: (2024)
by: Zhang, Jiashu, et al.
Published: (2024)
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
by: Jung, Hyunji, et al.
Published: (2026)
by: Jung, Hyunji, et al.
Published: (2026)
ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism
by: Aboudib, Alan, et al.
Published: (2026)
by: Aboudib, Alan, et al.
Published: (2026)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024)
by: Jeon, Byungsoo, et al.
Published: (2024)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
by: Wu, Houming, et al.
Published: (2024)
by: Wu, Houming, et al.
Published: (2024)
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training
by: Wu, Houming, et al.
Published: (2025)
by: Wu, Houming, et al.
Published: (2025)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026)
by: Lin, Jun-Liang, et al.
Published: (2026)
Delayed Random Partial Gradient Averaging for Federated Learning
by: Hu, Xinyi
Published: (2024)
by: Hu, Xinyi
Published: (2024)
Context Parallelism for Scalable Million-Token Inference
by: Yang, Amy, et al.
Published: (2024)
by: Yang, Amy, et al.
Published: (2024)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
by: Liu, Yangyijian, et al.
Published: (2025)
by: Liu, Yangyijian, et al.
Published: (2025)
TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
by: Shi, Ziji, et al.
Published: (2023)
by: Shi, Ziji, et al.
Published: (2023)
Parallel Split Learning with Global Sampling
by: Kohankhaki, Mohammad, et al.
Published: (2024)
by: Kohankhaki, Mohammad, et al.
Published: (2024)
Sampling Parallelism for Fast and Efficient Bayesian Learning
by: Özdemir, Asena Karolin, et al.
Published: (2026)
by: Özdemir, Asena Karolin, et al.
Published: (2026)
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
by: Han, Xueyuan, et al.
Published: (2024)
by: Han, Xueyuan, et al.
Published: (2024)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
by: Yang, Weihao, et al.
Published: (2025)
by: Yang, Weihao, et al.
Published: (2025)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
by: Hu, Xinyi, et al.
Published: (2026)
by: Hu, Xinyi, et al.
Published: (2026)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
by: Zeng, Zhichen, et al.
Published: (2026)
by: Zeng, Zhichen, et al.
Published: (2026)
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
by: Wagenländer, Marcel, et al.
Published: (2023)
by: Wagenländer, Marcel, et al.
Published: (2023)
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
by: Seal, Sudip K., et al.
Published: (2025)
by: Seal, Sudip K., et al.
Published: (2025)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs
by: Park, Chansung, et al.
Published: (2024)
by: Park, Chansung, et al.
Published: (2024)
Parallel-friendly Spatio-Temporal Graph Learning for Photovoltaic Degradation Analysis at Scale
by: Fan, Yangxin, et al.
Published: (2024)
by: Fan, Yangxin, et al.
Published: (2024)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
by: Kim, Han-Byul, et al.
Published: (2025)
by: Kim, Han-Byul, et al.
Published: (2025)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
by: Liu, Zhihong, et al.
Published: (2024)
by: Liu, Zhihong, et al.
Published: (2024)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
Online Parallel Multi-Task Relationship Learning via Alternating Direction Method of Multipliers
by: Li, Ruiyu, et al.
Published: (2024)
by: Li, Ruiyu, et al.
Published: (2024)
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2024)
by: Arfeen, Daiyaan, et al.
Published: (2024)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025)
by: Ranjan, Aditya K., et al.
Published: (2025)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
by: Wei, Cunyang, et al.
Published: (2026)
by: Wei, Cunyang, et al.
Published: (2026)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
by: Chen, Yanxi, et al.
Published: (2023)
by: Chen, Yanxi, et al.
Published: (2023)
Similar Items
-
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
by: Wan, Xinyi, et al.
Published: (2025) -
Pipeline Parallelism with Controllable Memory
by: Qi, Penghui, et al.
Published: (2024) -
Revisiting Parameter Server in LLM Post-Training
by: Wan, Xinyi, et al.
Published: (2026) -
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024) -
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
by: Guo, Jihu, et al.
Published: (2025)