MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912385805058048 |
|---|---|
| author | Jiang, Yinsicheng Fu, Yao Huang, Yeqi Nie, Ping Lu, Zhan Xue, Leyang He, Congjie Sit, Man-Kit Xue, Jilong Dong, Li Miao, Ziming Du, Dayou Xu, Tairan Zou, Kai Ponti, Edoardo Mai, Luo |
| author_facet | Jiang, Yinsicheng Fu, Yao Huang, Yeqi Nie, Ping Lu, Zhan Xue, Leyang He, Congjie Sit, Man-Kit Xue, Jilong Dong, Li Miao, Ziming Du, Dayou Xu, Tairan Zou, Kai Ponti, Edoardo Mai, Luo |
| contents | The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third-a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics-Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)-to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_11415 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems Jiang, Yinsicheng Fu, Yao Huang, Yeqi Nie, Ping Lu, Zhan Xue, Leyang He, Congjie Sit, Man-Kit Xue, Jilong Dong, Li Miao, Ziming Du, Dayou Xu, Tairan Zou, Kai Ponti, Edoardo Mai, Luo Machine Learning Distributed, Parallel, and Cluster Computing The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third-a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics-Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)-to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios. |
| title | MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2505.11415 |