Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Ziming, Tian, Boyu, Wang, Guoteng, Jiang, Zhen, Sun, Peng, Han, Zhenhua, Tang, Tian, Hu, Xiaohe, Jia, Yanmin, Zhang, Yan, Liu, He, Zhang, Mingjun, Zhang, Yiqi, Chen, Qiaoling, Cheng, Shenggan, Gao, Mingyu, You, Yang, Feng, Siyuan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912599569858560
author Liu, Ziming
Tian, Boyu
Wang, Guoteng
Jiang, Zhen
Sun, Peng
Han, Zhenhua
Tang, Tian
Hu, Xiaohe
Jia, Yanmin
Zhang, Yan
Liu, He
Zhang, Mingjun
Zhang, Yiqi
Chen, Qiaoling
Cheng, Shenggan
Gao, Mingyu
You, Yang
Feng, Siyuan
author_facet Liu, Ziming
Tian, Boyu
Wang, Guoteng
Jiang, Zhen
Sun, Peng
Han, Zhenhua
Tang, Tian
Hu, Xiaohe
Jia, Yanmin
Zhang, Yan
Liu, He
Zhang, Mingjun
Zhang, Yiqi
Chen, Qiaoling
Cheng, Shenggan
Gao, Mingyu
You, Yang
Feng, Siyuan
contents Mixture-of-Experts (MoE) models challenge serving infrastructures with dynamic, sparse expert utilization, causing instability on conventional systems designed for dense architectures. We propose EaaS, a novel serving system to enable efficient, scalable, and robust MoE deployment. Our system disaggregates MoE modules into independent, stateless services. This design enables fine-grained resource scaling and provides inherent fault tolerance by decoupling compute units. The architecture is powered by a high-performance, CPU-free peer-to-peer communication library that ensures minimal overhead and high throughput. Experiments confirm EaaS's scalability and efficiency, achieving performance comparable to monolithic systems while providing robust fault tolerance and strong scalability. EaaS incurs less than a 2% throughput reduction under simulated hardware failures that would otherwise halt monolithic architectures. It further saves up to 37.5% of computing resources through dynamic fine-grained adaptation to serving traffic, demonstrating strong resilience for large-scale MoE deployment in production.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
Liu, Ziming
Tian, Boyu
Wang, Guoteng
Jiang, Zhen
Sun, Peng
Han, Zhenhua
Tang, Tian
Hu, Xiaohe
Jia, Yanmin
Zhang, Yan
Liu, He
Zhang, Mingjun
Zhang, Yiqi
Chen, Qiaoling
Cheng, Shenggan
Gao, Mingyu
You, Yang
Feng, Siyuan
Distributed, Parallel, and Cluster Computing
Mixture-of-Experts (MoE) models challenge serving infrastructures with dynamic, sparse expert utilization, causing instability on conventional systems designed for dense architectures. We propose EaaS, a novel serving system to enable efficient, scalable, and robust MoE deployment. Our system disaggregates MoE modules into independent, stateless services. This design enables fine-grained resource scaling and provides inherent fault tolerance by decoupling compute units. The architecture is powered by a high-performance, CPU-free peer-to-peer communication library that ensures minimal overhead and high throughput. Experiments confirm EaaS's scalability and efficiency, achieving performance comparable to monolithic systems while providing robust fault tolerance and strong scalability. EaaS incurs less than a 2% throughput reduction under simulated hardware failures that would otherwise halt monolithic architectures. It further saves up to 37.5% of computing resources through dynamic fine-grained adaptation to serving traffic, demonstrating strong resilience for large-scale MoE deployment in production.
title Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.17863