Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Xinrui, Zhang, Hongxing, Zeng, Fanyi, Wei, Yongxian, Wang, Yizhi, Ling, Xitong, Li, Guanghao, Yuan, Chun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911074541895680
author Chen, Xinrui
Zhang, Hongxing
Zeng, Fanyi
Wei, Yongxian
Wang, Yizhi
Ling, Xitong
Li, Guanghao
Yuan, Chun
author_facet Chen, Xinrui
Zhang, Hongxing
Zeng, Fanyi
Wei, Yongxian
Wang, Yizhi
Ling, Xitong
Li, Guanghao
Yuan, Chun
contents Layer pruning has emerged as a promising technique for compressing large language models (LLMs) while achieving acceleration proportional to the pruning ratio. In this work, we identify that removing any layer induces a significant magnitude gap in hidden states, resulting in substantial performance degradation. To address this issue, we propose Prune&Comp, a novel plug-and-play layer pruning scheme that leverages magnitude compensation to mitigate such gaps in a training-free manner. Specifically, we first estimate the magnitude gap caused by layer removal and then eliminate this gap by rescaling the remaining weights offline, with zero runtime overhead incurred. We further demonstrate the advantages of Prune&Comp through an iterative pruning strategy. When integrated with an iterative prune-and-compensate loop, Prune&Comp consistently enhances existing layer pruning metrics. For instance, when 5 layers of LLaMA-3-8B are pruned using the prevalent block influence metric, Prune&Comp nearly halves the perplexity and retains 93.19\% of the original model's question-answering performance, outperforming the baseline by 4.01%.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
Chen, Xinrui
Zhang, Hongxing
Zeng, Fanyi
Wei, Yongxian
Wang, Yizhi
Ling, Xitong
Li, Guanghao
Yuan, Chun
Computation and Language
Layer pruning has emerged as a promising technique for compressing large language models (LLMs) while achieving acceleration proportional to the pruning ratio. In this work, we identify that removing any layer induces a significant magnitude gap in hidden states, resulting in substantial performance degradation. To address this issue, we propose Prune&Comp, a novel plug-and-play layer pruning scheme that leverages magnitude compensation to mitigate such gaps in a training-free manner. Specifically, we first estimate the magnitude gap caused by layer removal and then eliminate this gap by rescaling the remaining weights offline, with zero runtime overhead incurred. We further demonstrate the advantages of Prune&Comp through an iterative pruning strategy. When integrated with an iterative prune-and-compensate loop, Prune&Comp consistently enhances existing layer pruning metrics. For instance, when 5 layers of LLaMA-3-8B are pruned using the prevalent block influence metric, Prune&Comp nearly halves the perplexity and retains 93.19\% of the original model's question-answering performance, outperforming the baseline by 4.01%.
title Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
topic Computation and Language
url https://arxiv.org/abs/2507.18212