MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
Fuente:
arXiv
Guardado en:
| Autores principales: | Go, Seokjin, Mahajan, Divya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
por: Go, Seokjin, et al.
Publicado: (2026)
por: Go, Seokjin, et al.
Publicado: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
por: Zhu, Ruidong, et al.
Publicado: (2025)
por: Zhu, Ruidong, et al.
Publicado: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
por: Zhao, Adrian, et al.
Publicado: (2026)
por: Zhao, Adrian, et al.
Publicado: (2026)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
por: Go, Seokjin, et al.
Publicado: (2025)
por: Go, Seokjin, et al.
Publicado: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
por: Du, Zhixu, et al.
Publicado: (2023)
por: Du, Zhixu, et al.
Publicado: (2023)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
por: Shahout, Rana, et al.
Publicado: (2025)
por: Shahout, Rana, et al.
Publicado: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
por: Imani, HamidReza, et al.
Publicado: (2025)
por: Imani, HamidReza, et al.
Publicado: (2025)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
por: Wang, Minghe, et al.
Publicado: (2026)
por: Wang, Minghe, et al.
Publicado: (2026)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
por: Lin, Wenxiang, et al.
Publicado: (2025)
por: Lin, Wenxiang, et al.
Publicado: (2025)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
por: Wang, Irene, et al.
Publicado: (2026)
por: Wang, Irene, et al.
Publicado: (2026)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
por: Wu, Yongji, et al.
Publicado: (2025)
por: Wu, Yongji, et al.
Publicado: (2025)
Scattered Mixture-of-Experts Implementation
por: Tan, Shawn, et al.
Publicado: (2024)
por: Tan, Shawn, et al.
Publicado: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
por: Jiang, Yinsicheng, et al.
Publicado: (2025)
por: Jiang, Yinsicheng, et al.
Publicado: (2025)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
por: Jiang, Yinsicheng, et al.
Publicado: (2024)
por: Jiang, Yinsicheng, et al.
Publicado: (2024)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
por: Tairin, Suraiya, et al.
Publicado: (2025)
por: Tairin, Suraiya, et al.
Publicado: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
por: Liu, Mengfan, et al.
Publicado: (2025)
por: Liu, Mengfan, et al.
Publicado: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
por: Jin, Chao, et al.
Publicado: (2025)
por: Jin, Chao, et al.
Publicado: (2025)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
por: Zhan, Ziwei, et al.
Publicado: (2024)
por: Zhan, Ziwei, et al.
Publicado: (2024)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
por: Skiadopoulos, Athinagoras, et al.
Publicado: (2025)
por: Skiadopoulos, Athinagoras, et al.
Publicado: (2025)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
por: Lee, Seonho, et al.
Publicado: (2025)
por: Lee, Seonho, et al.
Publicado: (2025)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
pFedMoE: Data-Level Personalization with Mixture of Experts for Model-Heterogeneous Personalized Federated Learning
por: Yi, Liping, et al.
Publicado: (2024)
por: Yi, Liping, et al.
Publicado: (2024)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
por: Yu, Yanpeng, et al.
Publicado: (2025)
por: Yu, Yanpeng, et al.
Publicado: (2025)
Integrated Hardware Architecture and Device Placement Search
por: Wang, Irene, et al.
Publicado: (2024)
por: Wang, Irene, et al.
Publicado: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
por: Wu, Yongji, et al.
Publicado: (2024)
por: Wu, Yongji, et al.
Publicado: (2024)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
por: Luo, Shuqing, et al.
Publicado: (2025)
por: Luo, Shuqing, et al.
Publicado: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
por: Lee, Gunjun, et al.
Publicado: (2025)
por: Lee, Gunjun, et al.
Publicado: (2025)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
por: Yuan, Yueming, et al.
Publicado: (2025)
por: Yuan, Yueming, et al.
Publicado: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
por: Ma, Songkai, et al.
Publicado: (2025)
por: Ma, Songkai, et al.
Publicado: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
por: Wu, Tian, et al.
Publicado: (2025)
por: Wu, Tian, et al.
Publicado: (2025)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
por: Liu, Jiacheng, et al.
Publicado: (2024)
por: Liu, Jiacheng, et al.
Publicado: (2024)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
por: Jiang, Chenyu, et al.
Publicado: (2024)
por: Jiang, Chenyu, et al.
Publicado: (2024)
EarthSight: A Distributed Framework for Low-Latency Satellite Intelligence
por: Erol, Ansel Kaplan, et al.
Publicado: (2025)
por: Erol, Ansel Kaplan, et al.
Publicado: (2025)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
por: Tang, Peng, et al.
Publicado: (2024)
por: Tang, Peng, et al.
Publicado: (2024)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
por: Yao, Jinghan, et al.
Publicado: (2024)
por: Yao, Jinghan, et al.
Publicado: (2024)
Ejemplares similares
-
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
por: Go, Seokjin, et al.
Publicado: (2026) -
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
por: Zhu, Ruidong, et al.
Publicado: (2025) -
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
por: Zhao, Adrian, et al.
Publicado: (2026) -
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
por: Go, Seokjin, et al.
Publicado: (2025) -
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
por: Du, Zhixu, et al.
Publicado: (2023)