Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Weilin, Jiang, Juyong, Qin, Le, Cui, Junwei, Kim, Sunghun, Huang, Jiayi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
von: Cai, Weilin, et al.
Veröffentlicht: (2025)
von: Cai, Weilin, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
Scattered Mixture-of-Experts Implementation
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
BlackMamba: Mixture of Experts for State-Space Models
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
Fed-GAME: Personalized Federated Learning with Graph Attention Mixture-of-Experts For Time-Series Forecasting
von: Li, Yi, et al.
Veröffentlicht: (2026)
von: Li, Yi, et al.
Veröffentlicht: (2026)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
von: Zhan, Ziwei, et al.
Veröffentlicht: (2024)
von: Zhan, Ziwei, et al.
Veröffentlicht: (2024)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
pFedMoE: Data-Level Personalization with Mixture of Experts for Model-Heterogeneous Personalized Federated Learning
von: Yi, Liping, et al.
Veröffentlicht: (2024)
von: Yi, Liping, et al.
Veröffentlicht: (2024)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
von: Cai, Weilin, et al.
Veröffentlicht: (2024) -
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
von: Jiang, Juyong, et al.
Veröffentlicht: (2026) -
DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
von: Cai, Weilin, et al.
Veröffentlicht: (2025) -
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025) -
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)