ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ai, Mengting, Wei, Tianxin, Chen, Yifan, Zeng, Zhichen, Zhao, Ritchie, Varatkar, Girish, Rouhani, Bita Darvish, Tang, Xianfeng, Tong, Hanghang, He, Jingrui
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916648040005632
author Ai, Mengting
Wei, Tianxin
Chen, Yifan
Zeng, Zhichen
Zhao, Ritchie
Varatkar, Girish
Rouhani, Bita Darvish
Tang, Xianfeng
Tong, Hanghang
He, Jingrui
author_facet Ai, Mengting
Wei, Tianxin
Chen, Yifan
Zeng, Zhichen
Zhao, Ritchie
Varatkar, Girish
Rouhani, Bita Darvish
Tang, Xianfeng
Tong, Hanghang
He, Jingrui
contents Mixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for each input token. The sparse structure, while allowing constant time costs, results in space inefficiency: we still need to load all the model parameters during inference. We introduce ResMoE, an innovative MoE approximation framework that utilizes Wasserstein barycenter to extract a common expert (barycenter expert) and approximate the residuals between this barycenter expert and the original ones. ResMoE enhances the space efficiency for inference of large-scale MoE Transformers in a one-shot and data-agnostic manner without retraining while maintaining minimal accuracy loss, thereby paving the way for broader accessibility to large language models. We demonstrate the effectiveness of ResMoE through extensive experiments on Switch Transformer, Mixtral, and DeepSeekMoE models. The results show that ResMoE can reduce the number of parameters in an expert by up to 75% while maintaining comparable performance. The code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/ResMoE.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06881
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
Ai, Mengting
Wei, Tianxin
Chen, Yifan
Zeng, Zhichen
Zhao, Ritchie
Varatkar, Girish
Rouhani, Bita Darvish
Tang, Xianfeng
Tong, Hanghang
He, Jingrui
Machine Learning
Mixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for each input token. The sparse structure, while allowing constant time costs, results in space inefficiency: we still need to load all the model parameters during inference. We introduce ResMoE, an innovative MoE approximation framework that utilizes Wasserstein barycenter to extract a common expert (barycenter expert) and approximate the residuals between this barycenter expert and the original ones. ResMoE enhances the space efficiency for inference of large-scale MoE Transformers in a one-shot and data-agnostic manner without retraining while maintaining minimal accuracy loss, thereby paving the way for broader accessibility to large language models. We demonstrate the effectiveness of ResMoE through extensive experiments on Switch Transformer, Mixtral, and DeepSeekMoE models. The results show that ResMoE can reduce the number of parameters in an expert by up to 75% while maintaining comparable performance. The code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/ResMoE.
title ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
topic Machine Learning
url https://arxiv.org/abs/2503.06881