SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liang, Xun, Wang, Hanyu, Lai, Huayi, Niu, Simin, Song, Shichao, Yang, Jiawei, Zhao, Jihao, Xiong, Feiyu, Tang, Bo, Li, Zhiyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909532831088640
author Liang, Xun
Wang, Hanyu
Lai, Huayi
Niu, Simin
Song, Shichao
Yang, Jiawei
Zhao, Jihao
Xiong, Feiyu
Tang, Bo
Li, Zhiyu
author_facet Liang, Xun
Wang, Hanyu
Lai, Huayi
Niu, Simin
Song, Shichao
Yang, Jiawei
Zhao, Jihao
Xiong, Feiyu
Tang, Bo
Li, Zhiyu
contents Large Language Models have achieved remarkable success across various natural language processing tasks, yet their high computational cost during inference remains a major bottleneck. This paper introduces Sparse Expert Activation Pruning (SEAP), a training-free pruning method that selectively retains task-relevant parameters to reduce inference overhead. Inspired by the clustering patterns of hidden states and activations in LLMs, SEAP identifies task-specific expert activation patterns and prunes the model while preserving task performance and enhancing computational efficiency. Experimental results demonstrate that SEAP significantly reduces computational overhead while maintaining competitive accuracy. Notably, at 50% pruning, SEAP surpasses both WandA and FLAP by over 20%, and at 20% pruning, it incurs only a 2.2% performance drop compared to the dense model. These findings highlight SEAP's scalability and effectiveness, making it a promising approach for optimizing large-scale LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_07605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models
Liang, Xun
Wang, Hanyu
Lai, Huayi
Niu, Simin
Song, Shichao
Yang, Jiawei
Zhao, Jihao
Xiong, Feiyu
Tang, Bo
Li, Zhiyu
Computation and Language
Large Language Models have achieved remarkable success across various natural language processing tasks, yet their high computational cost during inference remains a major bottleneck. This paper introduces Sparse Expert Activation Pruning (SEAP), a training-free pruning method that selectively retains task-relevant parameters to reduce inference overhead. Inspired by the clustering patterns of hidden states and activations in LLMs, SEAP identifies task-specific expert activation patterns and prunes the model while preserving task performance and enhancing computational efficiency. Experimental results demonstrate that SEAP significantly reduces computational overhead while maintaining competitive accuracy. Notably, at 50% pruning, SEAP surpasses both WandA and FLAP by over 20%, and at 20% pruning, it incurs only a 2.2% performance drop compared to the dense model. These findings highlight SEAP's scalability and effectiveness, making it a promising approach for optimizing large-scale LLMs.
title SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2503.07605