QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Changhai, Zhou, Yuhua, Han, Shijie, Qiao, Qian, Li, Hongguang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917870056767488
author Zhou, Changhai
Zhou, Yuhua
Han, Shijie
Qiao, Qian
Li, Hongguang
author_facet Zhou, Changhai
Zhou, Yuhua
Han, Shijie
Qiao, Qian
Li, Hongguang
contents The rise of large language models (LLMs) has significantly advanced various natural language processing (NLP) tasks. However, the resource demands of these models pose substantial challenges. Structured pruning is an effective approach to reducing model size, but it often results in significant accuracy degradation, necessitating parameter updates to adapt. Unfortunately, such fine-tuning requires substantial memory, which limits its applicability. To address these challenges, we introduce quantization into the structured pruning framework to reduce memory consumption during both fine-tuning and inference. However, the combined errors from pruning and quantization increase the difficulty of fine-tuning, requiring a more refined quantization scheme. To this end, we propose QPruner, a novel framework that employs structured pruning to reduce model size, followed by a layer-wise mixed-precision quantization scheme. Quantization precisions are assigned to each layer based on their importance to the target task, and Bayesian optimization is employed to refine precision allocation strategies, ensuring a balance between model accuracy and memory efficiency. Extensive experiments on benchmark datasets demonstrate that QPruner significantly outperforms existing methods in memory savings while maintaining or improving model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11629
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language Models
Zhou, Changhai
Zhou, Yuhua
Han, Shijie
Qiao, Qian
Li, Hongguang
Machine Learning
The rise of large language models (LLMs) has significantly advanced various natural language processing (NLP) tasks. However, the resource demands of these models pose substantial challenges. Structured pruning is an effective approach to reducing model size, but it often results in significant accuracy degradation, necessitating parameter updates to adapt. Unfortunately, such fine-tuning requires substantial memory, which limits its applicability. To address these challenges, we introduce quantization into the structured pruning framework to reduce memory consumption during both fine-tuning and inference. However, the combined errors from pruning and quantization increase the difficulty of fine-tuning, requiring a more refined quantization scheme. To this end, we propose QPruner, a novel framework that employs structured pruning to reduce model size, followed by a layer-wise mixed-precision quantization scheme. Quantization precisions are assigned to each layer based on their importance to the target task, and Bayesian optimization is employed to refine precision allocation strategies, ensuring a balance between model accuracy and memory efficiency. Extensive experiments on benchmark datasets demonstrate that QPruner significantly outperforms existing methods in memory savings while maintaining or improving model performance.
title QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2412.11629