MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yinsicheng, Fu, Yao, Huang, Yeqi, Nie, Ping, Lu, Zhan, Xue, Leyang, He, Congjie, Sit, Man-Kit, Xue, Jilong, Dong, Li, Miao, Ziming, Du, Dayou, Xu, Tairan, Zou, Kai, Ponti, Edoardo, Mai, Luo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912385805058048
author Jiang, Yinsicheng
Fu, Yao
Huang, Yeqi
Nie, Ping
Lu, Zhan
Xue, Leyang
He, Congjie
Sit, Man-Kit
Xue, Jilong
Dong, Li
Miao, Ziming
Du, Dayou
Xu, Tairan
Zou, Kai
Ponti, Edoardo
Mai, Luo
author_facet Jiang, Yinsicheng
Fu, Yao
Huang, Yeqi
Nie, Ping
Lu, Zhan
Xue, Leyang
He, Congjie
Sit, Man-Kit
Xue, Jilong
Dong, Li
Miao, Ziming
Du, Dayou
Xu, Tairan
Zou, Kai
Ponti, Edoardo
Mai, Luo
contents The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third-a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics-Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)-to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11415
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
Jiang, Yinsicheng
Fu, Yao
Huang, Yeqi
Nie, Ping
Lu, Zhan
Xue, Leyang
He, Congjie
Sit, Man-Kit
Xue, Jilong
Dong, Li
Miao, Ziming
Du, Dayou
Xu, Tairan
Zou, Kai
Ponti, Edoardo
Mai, Luo
Machine Learning
Distributed, Parallel, and Cluster Computing
The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third-a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics-Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)-to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios.
title MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2505.11415