Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dash, Sajal, Wang, Feiyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
von: Seal, Sudip K., et al.
Veröffentlicht: (2025)
von: Seal, Sudip K., et al.
Veröffentlicht: (2025)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
Scalable Artificial Intelligence for Science: Perspectives, Methods and Exemplars
von: Brewer, Wesley, et al.
Veröffentlicht: (2024)
von: Brewer, Wesley, et al.
Veröffentlicht: (2024)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2024)
von: Wu, Houming, et al.
Veröffentlicht: (2024)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2025)
von: Wu, Houming, et al.
Veröffentlicht: (2025)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
von: Yu, Zhongkai, et al.
Veröffentlicht: (2025)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training
von: Brewer, Wesley, et al.
Veröffentlicht: (2025)
von: Brewer, Wesley, et al.
Veröffentlicht: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
PithTrain: A Compact and Agent-Native MoE Training System
von: Lai, Ruihang, et al.
Veröffentlicht: (2026)
von: Lai, Ruihang, et al.
Veröffentlicht: (2026)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024)
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024)
Zero Bubble Pipeline Parallelism
von: Qi, Penghui, et al.
Veröffentlicht: (2023)
von: Qi, Penghui, et al.
Veröffentlicht: (2023)
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
von: Jung, Hyunji, et al.
Veröffentlicht: (2026)
von: Jung, Hyunji, et al.
Veröffentlicht: (2026)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
von: Han, Xueyuan, et al.
Veröffentlicht: (2024)
von: Han, Xueyuan, et al.
Veröffentlicht: (2024)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism
von: Aboudib, Alan, et al.
Veröffentlicht: (2026)
von: Aboudib, Alan, et al.
Veröffentlicht: (2026)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
von: Chen, Ling, et al.
Veröffentlicht: (2026)
von: Chen, Ling, et al.
Veröffentlicht: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
von: Lin, Jun-Liang, et al.
Veröffentlicht: (2026)
von: Lin, Jun-Liang, et al.
Veröffentlicht: (2026)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025) -
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
von: Seal, Sudip K., et al.
Veröffentlicht: (2025) -
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026) -
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025) -
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025)