Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Wei, Lai, Zhiquan, Li, Shengwei, Liu, Weijie, Ge, Keshi, Shen, Ao, Su, Huayou, Li, Dongsheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
por: Li, Shengwei, et al.
Publicado: (2023)
por: Li, Shengwei, et al.
Publicado: (2023)
Fine-grained MoE Load Balancing with Linear Programming
por: Zhao, Chenqi, et al.
Publicado: (2025)
por: Zhao, Chenqi, et al.
Publicado: (2025)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
por: Qi, Shuyao, et al.
Publicado: (2026)
por: Qi, Shuyao, et al.
Publicado: (2026)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
por: Zheng, Size, et al.
Publicado: (2026)
por: Zheng, Size, et al.
Publicado: (2026)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
por: Wang, Yingping, et al.
Publicado: (2026)
por: Wang, Yingping, et al.
Publicado: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
por: Wang, Liujianfu, et al.
Publicado: (2025)
por: Wang, Liujianfu, et al.
Publicado: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
por: Nie, Xiaonan, et al.
Publicado: (2024)
por: Nie, Xiaonan, et al.
Publicado: (2024)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
por: Liu, Dennis, et al.
Publicado: (2025)
por: Liu, Dennis, et al.
Publicado: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
por: Ai, Xin, et al.
Publicado: (2024)
por: Ai, Xin, et al.
Publicado: (2024)
Accelerating Distributed MoE Training and Inference with Lina
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
Load Balanced Parallel Node Generation for Meshless Numerical Methods
por: Vehovar, Jon, et al.
Publicado: (2026)
por: Vehovar, Jon, et al.
Publicado: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
por: Luo, Jiajun, et al.
Publicado: (2024)
por: Luo, Jiajun, et al.
Publicado: (2024)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024)
por: Luo, Shuqing, et al.
Publicado: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
Sparse Checkpointing for Fast and Reliable MoE Training
por: Gandhi, Swapnil, et al.
Publicado: (2024)
por: Gandhi, Swapnil, et al.
Publicado: (2024)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
por: Sun, Xun, et al.
Publicado: (2026)
por: Sun, Xun, et al.
Publicado: (2026)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
por: Zhao, Lu, et al.
Publicado: (2025)
por: Zhao, Lu, et al.
Publicado: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
por: Zhang, Zheng, et al.
Publicado: (2025)
por: Zhang, Zheng, et al.
Publicado: (2025)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
por: Luo, Shuqing, et al.
Publicado: (2025)
por: Luo, Shuqing, et al.
Publicado: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
por: Li, Haoyang, et al.
Publicado: (2024)
por: Li, Haoyang, et al.
Publicado: (2024)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
por: Chen, Liangkun, et al.
Publicado: (2025)
por: Chen, Liangkun, et al.
Publicado: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
por: Han, Yu, et al.
Publicado: (2025)
por: Han, Yu, et al.
Publicado: (2025)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
por: Zhu, Qianchao, et al.
Publicado: (2026)
por: Zhu, Qianchao, et al.
Publicado: (2026)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
por: Tang, Xinru, et al.
Publicado: (2025)
por: Tang, Xinru, et al.
Publicado: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
por: Liu, Ziming, et al.
Publicado: (2025)
por: Liu, Ziming, et al.
Publicado: (2025)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
por: Song, Xiaoniu, et al.
Publicado: (2024)
por: Song, Xiaoniu, et al.
Publicado: (2024)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
por: Zhou, Bowen, et al.
Publicado: (2026)
por: Zhou, Bowen, et al.
Publicado: (2026)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
por: Chen, Chang, et al.
Publicado: (2025)
por: Chen, Chang, et al.
Publicado: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
por: Liu, Guowei, et al.
Publicado: (2026)
por: Liu, Guowei, et al.
Publicado: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
por: Qian, Yulei, et al.
Publicado: (2024)
por: Qian, Yulei, et al.
Publicado: (2024)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
por: Doucet, Zachary, et al.
Publicado: (2025)
por: Doucet, Zachary, et al.
Publicado: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
por: Liu, Di, et al.
Publicado: (2026)
por: Liu, Di, et al.
Publicado: (2026)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
por: Bu, Tianci, et al.
Publicado: (2026)
por: Bu, Tianci, et al.
Publicado: (2026)
Balancing Pipeline Parallelism with Vocabulary Parallelism
por: Yeung, Man Tsung, et al.
Publicado: (2024)
por: Yeung, Man Tsung, et al.
Publicado: (2024)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
por: Zeng, Zhichen, et al.
Publicado: (2026)
por: Zeng, Zhichen, et al.
Publicado: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
por: Pan, Xinglin, et al.
Publicado: (2025)
por: Pan, Xinglin, et al.
Publicado: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
por: Huang, En-Ming, et al.
Publicado: (2025)
por: Huang, En-Ming, et al.
Publicado: (2025)
When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
por: Zhu, Weihao, et al.
Publicado: (2025)
por: Zhu, Weihao, et al.
Publicado: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
por: Li, Haley, et al.
Publicado: (2026)
por: Li, Haley, et al.
Publicado: (2026)
Ejemplares similares
-
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
por: Li, Shengwei, et al.
Publicado: (2023) -
Fine-grained MoE Load Balancing with Linear Programming
por: Zhao, Chenqi, et al.
Publicado: (2025) -
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
por: Qi, Shuyao, et al.
Publicado: (2026) -
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
por: Zheng, Size, et al.
Publicado: (2026) -
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
por: Wang, Yingping, et al.
Publicado: (2026)