LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chitty-Venkata, Krishna Teja, Madireddy, Sandeep, Emani, Murali, Vishwanath, Venkatram
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914018518630400
author Chitty-Venkata, Krishna Teja
Madireddy, Sandeep
Emani, Murali
Vishwanath, Venkatram
author_facet Chitty-Venkata, Krishna Teja
Madireddy, Sandeep
Emani, Murali
Vishwanath, Venkatram
contents Mixture-of-Experts (MoE) models scale efficiently by activating only a subset of experts per token, offering a computationally sparse alternative to dense architectures. While prior post-training optimizations, such as inter- and intra-expert pruning, reduce memory usage they provide limited gains in inference-time compute efficiency. Moreover, existing MoE architectures typically activate a fixed number of experts uniformly across all layers, resulting in redundant computation and suboptimal performance. In this work, we first demonstrate that MoE pruning strategies improve only the memory footprint but do not significantly improve inference performance on GPU using optimized frameworks such as vLLM. To address this, we introduce LExI, a data-free optimization technique that determines the optimal number of active experts per layer in a pretrained MoE model. LExI leverages only the model weights to estimate the relative importance of each layer and adaptively assigns the number of active experts accordingly per layer. Experiments on state-of-the-art language and vision MoE benchmarks demonstrate that LExI significantly outperforms traditional MoE pruning approaches in terms of inference efficiency with negligible accuracy loss. For example, using LExI, Qwen1.5-MoE achieves the same throughput on Nvidia H100 GPU with 10% better accuracy than traditional expert pruning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
Chitty-Venkata, Krishna Teja
Madireddy, Sandeep
Emani, Murali
Vishwanath, Venkatram
Machine Learning
Mixture-of-Experts (MoE) models scale efficiently by activating only a subset of experts per token, offering a computationally sparse alternative to dense architectures. While prior post-training optimizations, such as inter- and intra-expert pruning, reduce memory usage they provide limited gains in inference-time compute efficiency. Moreover, existing MoE architectures typically activate a fixed number of experts uniformly across all layers, resulting in redundant computation and suboptimal performance. In this work, we first demonstrate that MoE pruning strategies improve only the memory footprint but do not significantly improve inference performance on GPU using optimized frameworks such as vLLM. To address this, we introduce LExI, a data-free optimization technique that determines the optimal number of active experts per layer in a pretrained MoE model. LExI leverages only the model weights to estimate the relative importance of each layer and adaptively assigns the number of active experts accordingly per layer. Experiments on state-of-the-art language and vision MoE benchmarks demonstrate that LExI significantly outperforms traditional MoE pruning approaches in terms of inference efficiency with negligible accuracy loss. For example, using LExI, Qwen1.5-MoE achieves the same throughput on Nvidia H100 GPU with 10% better accuracy than traditional expert pruning.
title LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
topic Machine Learning
url https://arxiv.org/abs/2509.02753