Entropy-Based Block Pruning for Efficient Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Liangwei, Xu, Yuhui, Tan, Juntao, Sahoo, Doyen, Savarese, Silvio, Xiong, Caiming, Wang, Huan, Heinecke, Shelby
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915229120593920
author Yang, Liangwei
Xu, Yuhui
Tan, Juntao
Sahoo, Doyen
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Heinecke, Shelby
author_facet Yang, Liangwei
Xu, Yuhui
Tan, Juntao
Sahoo, Doyen
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Heinecke, Shelby
contents As large language models continue to scale, their growing computational and storage demands pose significant challenges for real-world deployment. In this work, we investigate redundancy within Transformer-based models and propose an entropy-based pruning strategy to enhance efficiency while maintaining performance. Empirical analysis reveals that the entropy of hidden representations decreases in the early blocks but progressively increases across most subsequent blocks. This trend suggests that entropy serves as a more effective measure of information richness within computation blocks. Unlike cosine similarity, which primarily captures geometric relationships, entropy directly quantifies uncertainty and information content, making it a more reliable criterion for pruning. Extensive experiments demonstrate that our entropy-based pruning approach surpasses cosine similarity-based methods in reducing model size while preserving accuracy, offering a promising direction for efficient model deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Entropy-Based Block Pruning for Efficient Large Language Models
Yang, Liangwei
Xu, Yuhui
Tan, Juntao
Sahoo, Doyen
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Heinecke, Shelby
Computation and Language
Artificial Intelligence
As large language models continue to scale, their growing computational and storage demands pose significant challenges for real-world deployment. In this work, we investigate redundancy within Transformer-based models and propose an entropy-based pruning strategy to enhance efficiency while maintaining performance. Empirical analysis reveals that the entropy of hidden representations decreases in the early blocks but progressively increases across most subsequent blocks. This trend suggests that entropy serves as a more effective measure of information richness within computation blocks. Unlike cosine similarity, which primarily captures geometric relationships, entropy directly quantifies uncertainty and information content, making it a more reliable criterion for pruning. Extensive experiments demonstrate that our entropy-based pruning approach surpasses cosine similarity-based methods in reducing model size while preserving accuracy, offering a promising direction for efficient model deployment.
title Entropy-Based Block Pruning for Efficient Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.03794