Mixture Compressor for Mixture-of-Experts LLMs Gains More

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wei, Liao, Yue, Liu, Jianhui, He, Ruifei, Tan, Haoru, Zhang, Shiming, Li, Hongsheng, Liu, Si, Qi, Xiaojuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910839923015680
author Huang, Wei
Liao, Yue
Liu, Jianhui
He, Ruifei
Tan, Haoru
Zhang, Shiming
Li, Hongsheng
Liu, Si
Qi, Xiaojuan
author_facet Huang, Wei
Liao, Yue
Liu, Jianhui
He, Ruifei
Tan, Haoru
Zhang, Shiming
Li, Hongsheng
Liu, Si
Qi, Xiaojuan
contents Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading latency; and 2) the current activated experts are redundant, as many tokens may only require a single expert. Motivated by these issues, we investigate the MoE-LLMs and make two key observations: a) different experts exhibit varying behaviors on activation reconstruction error, routing scores, and activated frequencies, highlighting their differing importance, and b) not all tokens are equally important -- only a small subset is critical. Building on these insights, we propose MC, a training-free Mixture-Compressor for MoE-LLMs, which leverages the significance of both experts and tokens to achieve an extreme compression. First, to mitigate storage and loading overheads, we introduce Pre-Loading Mixed-Precision Quantization, which formulates the adaptive bit-width allocation as a Linear Programming problem, where the objective function balances multi-factors reflecting the importance of each expert. Additionally, we develop Online Dynamic Pruning, which identifies important tokens to retain and dynamically select activated experts for other tokens during inference to optimize efficiency while maintaining performance. Our MC integrates static quantization and dynamic pruning to collaboratively achieve extreme compression for MoE-LLMs with less accuracy loss, ensuring an optimal trade-off between performance and efficiency. Extensive experiments confirm the effectiveness of our approach. For instance, at 2.54 bits, MC compresses 76.6% of the model, with only a 3.8% average accuracy loss. During dynamic inference, we further reduce activated parameters by 15%, with a performance drop of less than 0.6%.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06270
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mixture Compressor for Mixture-of-Experts LLMs Gains More
Huang, Wei
Liao, Yue
Liu, Jianhui
He, Ruifei
Tan, Haoru
Zhang, Shiming
Li, Hongsheng
Liu, Si
Qi, Xiaojuan
Machine Learning
Computation and Language
Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading latency; and 2) the current activated experts are redundant, as many tokens may only require a single expert. Motivated by these issues, we investigate the MoE-LLMs and make two key observations: a) different experts exhibit varying behaviors on activation reconstruction error, routing scores, and activated frequencies, highlighting their differing importance, and b) not all tokens are equally important -- only a small subset is critical. Building on these insights, we propose MC, a training-free Mixture-Compressor for MoE-LLMs, which leverages the significance of both experts and tokens to achieve an extreme compression. First, to mitigate storage and loading overheads, we introduce Pre-Loading Mixed-Precision Quantization, which formulates the adaptive bit-width allocation as a Linear Programming problem, where the objective function balances multi-factors reflecting the importance of each expert. Additionally, we develop Online Dynamic Pruning, which identifies important tokens to retain and dynamically select activated experts for other tokens during inference to optimize efficiency while maintaining performance. Our MC integrates static quantization and dynamic pruning to collaboratively achieve extreme compression for MoE-LLMs with less accuracy loss, ensuring an optimal trade-off between performance and efficiency. Extensive experiments confirm the effectiveness of our approach. For instance, at 2.54 bits, MC compresses 76.6% of the model, with only a 3.8% average accuracy loss. During dynamic inference, we further reduce activated parameters by 15%, with a performance drop of less than 0.6%.
title Mixture Compressor for Mixture-of-Experts LLMs Gains More
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.06270