Scattered Mixture-of-Experts Implementation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Shawn, Shen, Yikang, Panda, Rameswar, Courville, Aaron |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
von: Zhan, Ziwei, et al.
Veröffentlicht: (2024)
von: Zhan, Ziwei, et al.
Veröffentlicht: (2024)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
Fed-GAME: Personalized Federated Learning with Graph Attention Mixture-of-Experts For Time-Series Forecasting
von: Li, Yi, et al.
Veröffentlicht: (2026)
von: Li, Yi, et al.
Veröffentlicht: (2026)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
pFedMoE: Data-Level Personalization with Mixture of Experts for Model-Heterogeneous Personalized Federated Learning
von: Yi, Liping, et al.
Veröffentlicht: (2024)
von: Yi, Liping, et al.
Veröffentlicht: (2024)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
Utility-Driven Speculative Decoding for Mixture-of-Experts
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
von: Sharma, Vyom, et al.
Veröffentlicht: (2026)
von: Sharma, Vyom, et al.
Veröffentlicht: (2026)
HDEE: Heterogeneous Domain Expert Ensemble
von: Ersoy, Oğuzhan, et al.
Veröffentlicht: (2025)
von: Ersoy, Oğuzhan, et al.
Veröffentlicht: (2025)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
Graph-Regularized Learning of Gaussian Mixture Models
von: Abdurakhmanova, Shamsiiat, et al.
Veröffentlicht: (2025)
von: Abdurakhmanova, Shamsiiat, et al.
Veröffentlicht: (2025)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
von: Vooturi, Dharma Teja, et al.
Veröffentlicht: (2026)
von: Vooturi, Dharma Teja, et al.
Veröffentlicht: (2026)
Improved Modelling of Federated Datasets using Mixtures-of-Dirichlet-Multinomials
von: Scott, Jonathan, et al.
Veröffentlicht: (2024)
von: Scott, Jonathan, et al.
Veröffentlicht: (2024)
Federated Learning for Misbehaviour Detection with Variational Autoencoders and Gaussian Mixture Models
von: Campos, Enrique Mármol, et al.
Veröffentlicht: (2024)
von: Campos, Enrique Mármol, et al.
Veröffentlicht: (2024)
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
FedGMI: Generative Model-Driven Federated Learning for Probabilistic Mixture Inference
von: Hou, Qijun, et al.
Veröffentlicht: (2026)
von: Hou, Qijun, et al.
Veröffentlicht: (2026)
FedDriveScore: Federated Scoring Driving Behavior with a Mixture of Metric Distributions
von: Lu, Lin
Veröffentlicht: (2024)
von: Lu, Lin
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024) -
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025) -
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025) -
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026) -
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)