Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jha, Saurav, Hashemzadeh, Maryam, Pasand, Ali Saheb, Parviz, Ali, Lee, Min-Joong, Knyazev, Boris
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2604.04356
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910105549668352
author Jha, Saurav
Hashemzadeh, Maryam
Pasand, Ali Saheb
Parviz, Ali
Lee, Min-Joong
Knyazev, Boris
author_facet Jha, Saurav
Hashemzadeh, Maryam
Pasand, Ali Saheb
Parviz, Ali
Lee, Min-Joong
Knyazev, Boris
contents Mixture-of-Experts (MoE) large language models (LLMs) are among the top-performing architectures. The largest models, often with hundreds of billions of parameters, pose significant memory challenges for deployment. Traditional approaches to reduce memory requirements include weight pruning and quantization. Motivated by the Router-weighted Expert Activation Pruning (REAP) that prunes experts, we propose a novel method, Router-weighted Expert Activation Merging (REAM). Instead of removing experts, REAM groups them and merges their weights, better preserving original performance. We evaluate REAM against REAP and other baselines across multiple MoE LLMs on diverse multiple-choice (MC) question answering and generative (GEN) benchmarks. Our results reveal a trade-off between MC and GEN performance that depends on the mix of calibration data. By controlling the mix of general, math and coding data, we examine the Pareto frontier of this trade-off and show that REAM often outperforms the baselines and in many cases is comparable to the original uncompressed models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04356
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle REAM: Merging Improves Pruning of Experts in LLMs
Jha, Saurav
Hashemzadeh, Maryam
Pasand, Ali Saheb
Parviz, Ali
Lee, Min-Joong
Knyazev, Boris
Artificial Intelligence
Computation and Language
Machine Learning
Performance
Mixture-of-Experts (MoE) large language models (LLMs) are among the top-performing architectures. The largest models, often with hundreds of billions of parameters, pose significant memory challenges for deployment. Traditional approaches to reduce memory requirements include weight pruning and quantization. Motivated by the Router-weighted Expert Activation Pruning (REAP) that prunes experts, we propose a novel method, Router-weighted Expert Activation Merging (REAM). Instead of removing experts, REAM groups them and merges their weights, better preserving original performance. We evaluate REAM against REAP and other baselines across multiple MoE LLMs on diverse multiple-choice (MC) question answering and generative (GEN) benchmarks. Our results reveal a trade-off between MC and GEN performance that depends on the mix of calibration data. By controlling the mix of general, math and coding data, we examine the Pareto frontier of this trade-off and show that REAM often outperforms the baselines and in many cases is comparable to the original uncompressed models.
title REAM: Merging Improves Pruning of Experts in LLMs
topic Artificial Intelligence
Computation and Language
Machine Learning
Performance
url https://arxiv.org/abs/2604.04356