Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912599569858560 |
|---|---|
| author | Liu, Ziming Tian, Boyu Wang, Guoteng Jiang, Zhen Sun, Peng Han, Zhenhua Tang, Tian Hu, Xiaohe Jia, Yanmin Zhang, Yan Liu, He Zhang, Mingjun Zhang, Yiqi Chen, Qiaoling Cheng, Shenggan Gao, Mingyu You, Yang Feng, Siyuan |
| author_facet | Liu, Ziming Tian, Boyu Wang, Guoteng Jiang, Zhen Sun, Peng Han, Zhenhua Tang, Tian Hu, Xiaohe Jia, Yanmin Zhang, Yan Liu, He Zhang, Mingjun Zhang, Yiqi Chen, Qiaoling Cheng, Shenggan Gao, Mingyu You, Yang Feng, Siyuan |
| contents | Mixture-of-Experts (MoE) models challenge serving infrastructures with dynamic, sparse expert utilization, causing instability on conventional systems designed for dense architectures. We propose EaaS, a novel serving system to enable efficient, scalable, and robust MoE deployment. Our system disaggregates MoE modules into independent, stateless services. This design enables fine-grained resource scaling and provides inherent fault tolerance by decoupling compute units. The architecture is powered by a high-performance, CPU-free peer-to-peer communication library that ensures minimal overhead and high throughput. Experiments confirm EaaS's scalability and efficiency, achieving performance comparable to monolithic systems while providing robust fault tolerance and strong scalability. EaaS incurs less than a 2% throughput reduction under simulated hardware failures that would otherwise halt monolithic architectures. It further saves up to 37.5% of computing resources through dynamic fine-grained adaptation to serving traffic, demonstrating strong resilience for large-scale MoE deployment in production. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17863 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving Liu, Ziming Tian, Boyu Wang, Guoteng Jiang, Zhen Sun, Peng Han, Zhenhua Tang, Tian Hu, Xiaohe Jia, Yanmin Zhang, Yan Liu, He Zhang, Mingjun Zhang, Yiqi Chen, Qiaoling Cheng, Shenggan Gao, Mingyu You, Yang Feng, Siyuan Distributed, Parallel, and Cluster Computing Mixture-of-Experts (MoE) models challenge serving infrastructures with dynamic, sparse expert utilization, causing instability on conventional systems designed for dense architectures. We propose EaaS, a novel serving system to enable efficient, scalable, and robust MoE deployment. Our system disaggregates MoE modules into independent, stateless services. This design enables fine-grained resource scaling and provides inherent fault tolerance by decoupling compute units. The architecture is powered by a high-performance, CPU-free peer-to-peer communication library that ensures minimal overhead and high throughput. Experiments confirm EaaS's scalability and efficiency, achieving performance comparable to monolithic systems while providing robust fault tolerance and strong scalability. EaaS incurs less than a 2% throughput reduction under simulated hardware failures that would otherwise halt monolithic architectures. It further saves up to 37.5% of computing resources through dynamic fine-grained adaptation to serving traffic, demonstrating strong resilience for large-scale MoE deployment in production. |
| title | Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2509.17863 |