Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Zeliang, Ghosh, Nikhil, Liu, Jiani, Yu, Bin, Liu, Xiaodong
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2604.06542
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910110304960512
author Zhang, Zeliang
Ghosh, Nikhil
Liu, Jiani
Yu, Bin
Liu, Xiaodong
author_facet Zhang, Zeliang
Ghosh, Nikhil
Liu, Jiani
Yu, Bin
Liu, Xiaodong
contents Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by activating only a subset of experts per forward pass, improving efficiency without sacrificing performance. However, the large number of expert parameters still leads to substantial memory consumption. Existing pruning methods typically allocate budgets uniformly across layers, overlooking the heterogeneous redundancy that arises in sparse MoEs. We propose GRAPE (Global Redundancy-Aware Pruning of Experts, a global pruning strategy that dynamically allocates pruning budgets based on cross-layer redundancy. Experiments on Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS show that, under the same pruning budget, GRAPE consistently achieves the best average performance. On the three main models reported in the paper, it improves average accuracy over the strongest local baseline by 1.40% on average across pruning settings, with gains of up to 2.45%.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06542
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Does a Global Perspective Help Prune Sparse MoEs Elegantly?
Zhang, Zeliang
Ghosh, Nikhil
Liu, Jiani
Yu, Bin
Liu, Xiaodong
Computation and Language
Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by activating only a subset of experts per forward pass, improving efficiency without sacrificing performance. However, the large number of expert parameters still leads to substantial memory consumption. Existing pruning methods typically allocate budgets uniformly across layers, overlooking the heterogeneous redundancy that arises in sparse MoEs. We propose GRAPE (Global Redundancy-Aware Pruning of Experts, a global pruning strategy that dynamically allocates pruning budgets based on cross-layer redundancy. Experiments on Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS show that, under the same pruning budget, GRAPE consistently achieves the best average performance. On the three main models reported in the paper, it improves average accuracy over the strongest local baseline by 1.40% on average across pruning settings, with gains of up to 2.45%.
title Does a Global Perspective Help Prune Sparse MoEs Elegantly?
topic Computation and Language
url https://arxiv.org/abs/2604.06542