ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Men, Xin, Xu, Mingyu, Zhang, Qingyu, Wang, Bingning, Lin, Hongyu, Lu, Yaojie, Han, Xianpei, Chen, Weipeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910644884733952
author Men, Xin
Xu, Mingyu
Zhang, Qingyu
Wang, Bingning
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Chen, Weipeng
author_facet Men, Xin
Xu, Mingyu
Zhang, Qingyu
Wang, Bingning
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Chen, Weipeng
contents As Large Language Models (LLMs) continue to advance in performance, their size has escalated significantly, with current LLMs containing billions or even trillions of parameters. However, in this study, we discovered that many layers of LLMs exhibit high similarity, and some layers play a negligible role in network functionality. Based on this observation, we define a metric called Block Influence (BI) to gauge the significance of each layer in LLMs. We then propose a straightforward pruning approach: layer removal, in which we directly delete the redundant layers in LLMs based on their BI scores. Experiments demonstrate that our method, which we call ShortGPT, significantly outperforms previous state-of-the-art (SOTA) methods in model pruning. Moreover, ShortGPT is orthogonal to quantization-like methods, enabling further reduction in parameters and computation. The ability to achieve better results through simple layer removal, as opposed to more complex pruning techniques, suggests a high degree of redundancy in the model architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03853
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Men, Xin
Xu, Mingyu
Zhang, Qingyu
Wang, Bingning
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Chen, Weipeng
Computation and Language
As Large Language Models (LLMs) continue to advance in performance, their size has escalated significantly, with current LLMs containing billions or even trillions of parameters. However, in this study, we discovered that many layers of LLMs exhibit high similarity, and some layers play a negligible role in network functionality. Based on this observation, we define a metric called Block Influence (BI) to gauge the significance of each layer in LLMs. We then propose a straightforward pruning approach: layer removal, in which we directly delete the redundant layers in LLMs based on their BI scores. Experiments demonstrate that our method, which we call ShortGPT, significantly outperforms previous state-of-the-art (SOTA) methods in model pruning. Moreover, ShortGPT is orthogonal to quantization-like methods, enabling further reduction in parameters and computation. The ability to achieve better results through simple layer removal, as opposed to more complex pruning techniques, suggests a high degree of redundancy in the model architecture.
title ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
topic Computation and Language
url https://arxiv.org/abs/2403.03853