FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Yunqi, Hu, Bing, Mashhadi, Mahdi Boloursaz, Jin, A-Long, Zhang, Yanfeng, Xiao, Pei, Tafazolli, Rahim, Debbah, Merouane |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
por: Han, Yunhe, et al.
Publicado: (2026)
por: Han, Yunhe, et al.
Publicado: (2026)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024)
por: Luo, Shuqing, et al.
Publicado: (2024)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
por: Qian, Yulei, et al.
Publicado: (2024)
por: Qian, Yulei, et al.
Publicado: (2024)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
por: Zhang, Han, et al.
Publicado: (2026)
por: Zhang, Han, et al.
Publicado: (2026)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
por: Yu, Dianhai, et al.
Publicado: (2022)
por: Yu, Dianhai, et al.
Publicado: (2022)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
por: Zhang, Zhexiang, et al.
Publicado: (2025)
por: Zhang, Zhexiang, et al.
Publicado: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
por: Chen, Tiancheng, et al.
Publicado: (2025)
por: Chen, Tiancheng, et al.
Publicado: (2025)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
por: Singh, Gursimran, et al.
Publicado: (2025)
por: Singh, Gursimran, et al.
Publicado: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
por: Wu, Yongji, et al.
Publicado: (2025)
por: Wu, Yongji, et al.
Publicado: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
por: Shen, Zixu, et al.
Publicado: (2025)
por: Shen, Zixu, et al.
Publicado: (2025)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
por: Yuan, Yueming, et al.
Publicado: (2025)
por: Yuan, Yueming, et al.
Publicado: (2025)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
por: Wang, Liujianfu, et al.
Publicado: (2025)
por: Wang, Liujianfu, et al.
Publicado: (2025)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
por: Wang, Zhixin, et al.
Publicado: (2025)
por: Wang, Zhixin, et al.
Publicado: (2025)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
por: Zhao, Lu, et al.
Publicado: (2025)
por: Zhao, Lu, et al.
Publicado: (2025)
Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
por: Wang, Xinbo, et al.
Publicado: (2025)
por: Wang, Xinbo, et al.
Publicado: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
por: Liu, Ziming, et al.
Publicado: (2025)
por: Liu, Ziming, et al.
Publicado: (2025)
Scalable Training of Mixture-of-Experts Models with Megatron Core
por: Yan, Zijie, et al.
Publicado: (2026)
por: Yan, Zijie, et al.
Publicado: (2026)
Accelerating Distributed MoE Training and Inference with Lina
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
por: Zheng, Size, et al.
Publicado: (2026)
por: Zheng, Size, et al.
Publicado: (2026)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
por: Jin, Chao, et al.
Publicado: (2025)
por: Jin, Chao, et al.
Publicado: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
por: Wu, Tian, et al.
Publicado: (2025)
por: Wu, Tian, et al.
Publicado: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
por: Du, Zhixu, et al.
Publicado: (2023)
por: Du, Zhixu, et al.
Publicado: (2023)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
por: Go, Seokjin, et al.
Publicado: (2025)
por: Go, Seokjin, et al.
Publicado: (2025)
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
por: Qin, Shengling, et al.
Publicado: (2025)
por: Qin, Shengling, et al.
Publicado: (2025)
DeFT: Mitigating Data Dependencies for Flexible Communication Scheduling in Distributed Training
por: Meng, Lin, et al.
Publicado: (2025)
por: Meng, Lin, et al.
Publicado: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
por: Wu, Yongji, et al.
Publicado: (2024)
por: Wu, Yongji, et al.
Publicado: (2024)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
por: Pan, Xinglin, et al.
Publicado: (2025)
por: Pan, Xinglin, et al.
Publicado: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
por: Zhang, Zheng, et al.
Publicado: (2025)
por: Zhang, Zheng, et al.
Publicado: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
Scalable HPC Job Scheduling and Resource Management in SST
por: Abdurahman, Abubeker, et al.
Publicado: (2025)
por: Abdurahman, Abubeker, et al.
Publicado: (2025)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
por: Sun, Yifan, et al.
Publicado: (2026)
por: Sun, Yifan, et al.
Publicado: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
por: Xu, Heng, et al.
Publicado: (2025)
por: Xu, Heng, et al.
Publicado: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
por: Hui, Xinning, et al.
Publicado: (2024)
por: Hui, Xinning, et al.
Publicado: (2024)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
por: Tang, Zhenheng, et al.
Publicado: (2025)
por: Tang, Zhenheng, et al.
Publicado: (2025)
Ejemplares similares
-
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
por: Han, Yunhe, et al.
Publicado: (2026) -
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
por: Yang, Zheming, et al.
Publicado: (2025) -
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025) -
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024) -
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
por: Qian, Yulei, et al.
Publicado: (2024)