Streamlining Redundant Layers to Compress Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Xiaodong, Hu, Yuxuan, Zhang, Jing, Wang, Yanling, Li, Cuiping, Chen, Hong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915120634920960
author Chen, Xiaodong
Hu, Yuxuan
Zhang, Jing
Wang, Yanling
Li, Cuiping
Chen, Hong
author_facet Chen, Xiaodong
Hu, Yuxuan
Zhang, Jing
Wang, Yanling
Li, Cuiping
Chen, Hong
contents This paper introduces LLM-Streamline, a pioneer work on layer pruning for large language models (LLMs). It is based on the observation that different layers have varying impacts on hidden states, enabling the identification of less important layers to be pruned.LLM-Streamline comprises two parts: layer pruning, which removes consecutive layers with the lowest importance based on target sparsity, and layer replacement, a novel module that trains a lightweight network to replace the pruned layers to mitigate performance loss. Additionally, a new metric called stability is proposed to address the limitations of the widely used accuracy metric in evaluating model compression. Experiments show that LLM-Streamline outperforms both previous and concurrent state-of-the-art pruning methods in terms of both performance and training efficiency.Our code is available at https://github.com/RUCKBReasoning/LLM-Streamline
format Preprint
id arxiv_https___arxiv_org_abs_2403_19135
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Streamlining Redundant Layers to Compress Large Language Models
Chen, Xiaodong
Hu, Yuxuan
Zhang, Jing
Wang, Yanling
Li, Cuiping
Chen, Hong
Computation and Language
Artificial Intelligence
I.2.7
This paper introduces LLM-Streamline, a pioneer work on layer pruning for large language models (LLMs). It is based on the observation that different layers have varying impacts on hidden states, enabling the identification of less important layers to be pruned.LLM-Streamline comprises two parts: layer pruning, which removes consecutive layers with the lowest importance based on target sparsity, and layer replacement, a novel module that trains a lightweight network to replace the pruned layers to mitigate performance loss. Additionally, a new metric called stability is proposed to address the limitations of the widely used accuracy metric in evaluating model compression. Experiments show that LLM-Streamline outperforms both previous and concurrent state-of-the-art pruning methods in terms of both performance and training efficiency.Our code is available at https://github.com/RUCKBReasoning/LLM-Streamline
title Streamlining Redundant Layers to Compress Large Language Models
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2403.19135