Streamlining Redundant Layers to Compress Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915120634920960 |
|---|---|
| author | Chen, Xiaodong Hu, Yuxuan Zhang, Jing Wang, Yanling Li, Cuiping Chen, Hong |
| author_facet | Chen, Xiaodong Hu, Yuxuan Zhang, Jing Wang, Yanling Li, Cuiping Chen, Hong |
| contents | This paper introduces LLM-Streamline, a pioneer work on layer pruning for large language models (LLMs). It is based on the observation that different layers have varying impacts on hidden states, enabling the identification of less important layers to be pruned.LLM-Streamline comprises two parts: layer pruning, which removes consecutive layers with the lowest importance based on target sparsity, and layer replacement, a novel module that trains a lightweight network to replace the pruned layers to mitigate performance loss. Additionally, a new metric called stability is proposed to address the limitations of the widely used accuracy metric in evaluating model compression. Experiments show that LLM-Streamline outperforms both previous and concurrent state-of-the-art pruning methods in terms of both performance and training efficiency.Our code is available at https://github.com/RUCKBReasoning/LLM-Streamline |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_19135 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Streamlining Redundant Layers to Compress Large Language Models Chen, Xiaodong Hu, Yuxuan Zhang, Jing Wang, Yanling Li, Cuiping Chen, Hong Computation and Language Artificial Intelligence I.2.7 This paper introduces LLM-Streamline, a pioneer work on layer pruning for large language models (LLMs). It is based on the observation that different layers have varying impacts on hidden states, enabling the identification of less important layers to be pruned.LLM-Streamline comprises two parts: layer pruning, which removes consecutive layers with the lowest importance based on target sparsity, and layer replacement, a novel module that trains a lightweight network to replace the pruned layers to mitigate performance loss. Additionally, a new metric called stability is proposed to address the limitations of the widely used accuracy metric in evaluating model compression. Experiments show that LLM-Streamline outperforms both previous and concurrent state-of-the-art pruning methods in terms of both performance and training efficiency.Our code is available at https://github.com/RUCKBReasoning/LLM-Streamline |
| title | Streamlining Redundant Layers to Compress Large Language Models |
| topic | Computation and Language Artificial Intelligence I.2.7 |
| url | https://arxiv.org/abs/2403.19135 |