CoSMoEs: Compact Sparse Mixture of Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huber, Patrick, Shrivastava, Akshat, Chang, Ernie, Sankar, Chinnadhurai, Aly, Ahmed, Sagar, Adithya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917941163851776
author Huber, Patrick
Shrivastava, Akshat
Chang, Ernie
Sankar, Chinnadhurai
Aly, Ahmed
Sagar, Adithya
author_facet Huber, Patrick
Shrivastava, Akshat
Chang, Ernie
Sankar, Chinnadhurai
Aly, Ahmed
Sagar, Adithya
contents Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device inference. Specifically, we tackle the three main on-device dimensions: Quality, Memory and Latency. Along the quality axis, we show that in a fair evaluation (removing confounding factors) MoE architectures outperform FLOP-aligned dense models at on-device scale. We introduce weight-decomposed experts, further improving the MoE model performance. Regarding model memory and latency, we significantly improve model offloading efficiency and, in turn, reduce model inference latency.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoSMoEs: Compact Sparse Mixture of Experts
Huber, Patrick
Shrivastava, Akshat
Chang, Ernie
Sankar, Chinnadhurai
Aly, Ahmed
Sagar, Adithya
Machine Learning
Computation and Language
Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device inference. Specifically, we tackle the three main on-device dimensions: Quality, Memory and Latency. Along the quality axis, we show that in a fair evaluation (removing confounding factors) MoE architectures outperform FLOP-aligned dense models at on-device scale. We introduce weight-decomposed experts, further improving the MoE model performance. Regarding model memory and latency, we significantly improve model offloading efficiency and, in turn, reduce model inference latency.
title CoSMoEs: Compact Sparse Mixture of Experts
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2503.00245