Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Mengfan, Wang, Wei, Wu, Chuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
von: Zhan, Ziwei, et al.
Veröffentlicht: (2024)
von: Zhan, Ziwei, et al.
Veröffentlicht: (2024)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
High-Performance Serverless Computing: A Systematic Literature Review on Serverless for HPC, AI, and Big Data
von: Besozzi, Valerio, et al.
Veröffentlicht: (2026)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2026)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
Scattered Mixture-of-Experts Implementation
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
von: Chen, Qian, et al.
Veröffentlicht: (2025)
von: Chen, Qian, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
pFedMoE: Data-Level Personalization with Mixture of Experts for Model-Heterogeneous Personalized Federated Learning
von: Yi, Liping, et al.
Veröffentlicht: (2024)
von: Yi, Liping, et al.
Veröffentlicht: (2024)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
FedGMI: Generative Model-Driven Federated Learning for Probabilistic Mixture Inference
von: Hou, Qijun, et al.
Veröffentlicht: (2026)
von: Hou, Qijun, et al.
Veröffentlicht: (2026)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
Fed-GAME: Personalized Federated Learning with Graph Attention Mixture-of-Experts For Time-Series Forecasting
von: Li, Yi, et al.
Veröffentlicht: (2026)
von: Li, Yi, et al.
Veröffentlicht: (2026)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
von: Vooturi, Dharma Teja, et al.
Veröffentlicht: (2026)
von: Vooturi, Dharma Teja, et al.
Veröffentlicht: (2026)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
Echo: Simulating Distributed Training At Scale
von: Feng, Yicheng, et al.
Veröffentlicht: (2024)
von: Feng, Yicheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
von: Wang, Minghe, et al.
Veröffentlicht: (2026) -
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024) -
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024) -
A Survey on Inference Optimization Techniques for Mixture of Experts Models
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024) -
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025)