CoSMoEs: Compact Sparse Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917941163851776 |
|---|---|
| author | Huber, Patrick Shrivastava, Akshat Chang, Ernie Sankar, Chinnadhurai Aly, Ahmed Sagar, Adithya |
| author_facet | Huber, Patrick Shrivastava, Akshat Chang, Ernie Sankar, Chinnadhurai Aly, Ahmed Sagar, Adithya |
| contents | Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device inference. Specifically, we tackle the three main on-device dimensions: Quality, Memory and Latency. Along the quality axis, we show that in a fair evaluation (removing confounding factors) MoE architectures outperform FLOP-aligned dense models at on-device scale. We introduce weight-decomposed experts, further improving the MoE model performance. Regarding model memory and latency, we significantly improve model offloading efficiency and, in turn, reduce model inference latency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_00245 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CoSMoEs: Compact Sparse Mixture of Experts Huber, Patrick Shrivastava, Akshat Chang, Ernie Sankar, Chinnadhurai Aly, Ahmed Sagar, Adithya Machine Learning Computation and Language Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device inference. Specifically, we tackle the three main on-device dimensions: Quality, Memory and Latency. Along the quality axis, we show that in a fair evaluation (removing confounding factors) MoE architectures outperform FLOP-aligned dense models at on-device scale. We introduce weight-decomposed experts, further improving the MoE model performance. Regarding model memory and latency, we significantly improve model offloading efficiency and, in turn, reduce model inference latency. |
| title | CoSMoEs: Compact Sparse Mixture of Experts |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2503.00245 |