PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arfeen, Daiyaan, Zhang, Zhen, Fu, Xinwei, Ganger, Gregory R., Wang, Yida |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024)
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
Efficient Training on Multiple Consumer GPUs with RoundPipe
von: Luo, Yibin, et al.
Veröffentlicht: (2026)
von: Luo, Yibin, et al.
Veröffentlicht: (2026)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
SkipPipe: Partial and Reordered Pipelining Framework for Training LLMs in Heterogeneous Networks
von: Blagoev, Nikolay, et al.
Veröffentlicht: (2025)
von: Blagoev, Nikolay, et al.
Veröffentlicht: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
TTrace: Lightweight Error Checking and Diagnosis for Distributed Training
von: Jiang, Haitian, et al.
Veröffentlicht: (2025)
von: Jiang, Haitian, et al.
Veröffentlicht: (2025)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
von: Bai, Xu, et al.
Veröffentlicht: (2026)
von: Bai, Xu, et al.
Veröffentlicht: (2026)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
von: Chen, Shawn Shuoshuo, et al.
Veröffentlicht: (2025)
von: Chen, Shawn Shuoshuo, et al.
Veröffentlicht: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
von: Chen, Tiancheng, et al.
Veröffentlicht: (2025)
von: Chen, Tiancheng, et al.
Veröffentlicht: (2025)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
von: Butler, Branden, et al.
Veröffentlicht: (2024)
von: Butler, Branden, et al.
Veröffentlicht: (2024)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2024)
von: Wu, Houming, et al.
Veröffentlicht: (2024)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Zero Bubble Pipeline Parallelism
von: Qi, Penghui, et al.
Veröffentlicht: (2023)
von: Qi, Penghui, et al.
Veröffentlicht: (2023)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
von: Li, Hongpei, et al.
Veröffentlicht: (2025)
von: Li, Hongpei, et al.
Veröffentlicht: (2025)
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
von: Jiang, Chenyu, et al.
Veröffentlicht: (2025)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2025)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2025)
von: Wu, Houming, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training
von: Barley, Daniel, et al.
Veröffentlicht: (2026)
von: Barley, Daniel, et al.
Veröffentlicht: (2026)
An inherently parallel H2-ULV factorization for solving dense linear systems on GPUs
von: Ma, Qianxiang, et al.
Veröffentlicht: (2025)
von: Ma, Qianxiang, et al.
Veröffentlicht: (2025)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025) -
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024) -
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024) -
Efficient Training on Multiple Consumer GPUs with RoundPipe
von: Luo, Yibin, et al.
Veröffentlicht: (2026) -
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)