LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917135530328064 |
|---|---|
| author | Su, Yupeng Guan, Ziyi Liu, Xiaoqun Jin, Tianlai Wu, Dongkuan Chen, Zhengfei Chesi, Graziano Wong, Ngai Yu, Hao |
| author_facet | Su, Yupeng Guan, Ziyi Liu, Xiaoqun Jin, Tianlai Wu, Dongkuan Chen, Zhengfei Chesi, Graziano Wong, Ngai Yu, Hao |
| contents | Large language models (LLMs) have seen substantial growth, necessitating efficient model pruning techniques. Existing post-training pruning methods primarily measure weight importance in converged dense models, often overlooking changes in weight significance during the pruning process, leading to performance degradation. To address this issue, we present LLM-Barber (Block-Aware Rebuilder for Sparsity Mask in One-Shot), a novel one-shot pruning framework that rebuilds the sparsity mask of pruned models without any retraining or weight reconstruction. LLM-Barber incorporates block-aware error optimization across Self-Attention and MLP blocks, facilitating global performance optimization. We are the first to employ the product of weights and gradients as a pruning metric in the context of LLM post-training pruning. This enables accurate identification of weight importance in massive models and significantly reduces computational complexity compared to methods using secondorder information. Our experiments show that LLM-Barber efficiently prunes models from LLaMA and OPT families (7B to 13B) on a single A100 GPU in just 30 minutes, achieving state-of-the-art results in both perplexity and zero-shot performance across various language benchmarks. Code is available at https://github.com/YupengSu/LLM-Barber. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_10631 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models Su, Yupeng Guan, Ziyi Liu, Xiaoqun Jin, Tianlai Wu, Dongkuan Chen, Zhengfei Chesi, Graziano Wong, Ngai Yu, Hao Machine Learning Artificial Intelligence Computation and Language Large language models (LLMs) have seen substantial growth, necessitating efficient model pruning techniques. Existing post-training pruning methods primarily measure weight importance in converged dense models, often overlooking changes in weight significance during the pruning process, leading to performance degradation. To address this issue, we present LLM-Barber (Block-Aware Rebuilder for Sparsity Mask in One-Shot), a novel one-shot pruning framework that rebuilds the sparsity mask of pruned models without any retraining or weight reconstruction. LLM-Barber incorporates block-aware error optimization across Self-Attention and MLP blocks, facilitating global performance optimization. We are the first to employ the product of weights and gradients as a pruning metric in the context of LLM post-training pruning. This enables accurate identification of weight importance in massive models and significantly reduces computational complexity compared to methods using secondorder information. Our experiments show that LLM-Barber efficiently prunes models from LLaMA and OPT families (7B to 13B) on a single A100 GPU in just 30 minutes, achieving state-of-the-art results in both perplexity and zero-shot performance across various language benchmarks. Code is available at https://github.com/YupengSu/LLM-Barber. |
| title | LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2408.10631 |