GWQ: Gradient-Aware Weight Quantization for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908383760613376 |
|---|---|
| author | Shao, Yihua Gu, Yan Chen, Siyu Liu, Haiyang Zhu, Zixian Ling, Zijian Yan, Minxi Yan, Ziyang Zhang, Chenyu Magno, Michele Qin, Haotong Wang, Yan Guo, Jingcai Shao, Ling Tang, Hao |
| author_facet | Shao, Yihua Gu, Yan Chen, Siyu Liu, Haiyang Zhu, Zixian Ling, Zijian Yan, Minxi Yan, Ziyang Zhang, Chenyu Magno, Michele Qin, Haotong Wang, Yan Guo, Jingcai Shao, Ling Tang, Hao |
| contents | Large language models (LLMs) show impressive performance in solving complex language tasks. However, its large number of parameters presents significant challenges for the deployment. So, compressing LLMs to low bits can enable to deploy on resource-constrained devices. To address this problem, we propose gradient-aware weight quantization (GWQ), the first quantization approach for low-bit weight quantization that leverages gradients to localize outliers, requiring only a minimal amount of calibration data for outlier detection. GWQ retains the top 1\% outliers preferentially at FP16 precision, while the remaining non-outlier weights are stored in a low-bit. We widely evaluate GWQ on different task include language modeling, grounding detection, massive multitask language understanding and vision-language question and answering. Results show that models quantified by GWQ performs better than other quantization method. During quantization process, GWQ only need one calibration set to realize effective quant. Also, GWQ achieves 1.2x inference speedup in comparison to the original model and effectively reduces the inference memory. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_00850 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GWQ: Gradient-Aware Weight Quantization for Large Language Models Shao, Yihua Gu, Yan Chen, Siyu Liu, Haiyang Zhu, Zixian Ling, Zijian Yan, Minxi Yan, Ziyang Zhang, Chenyu Magno, Michele Qin, Haotong Wang, Yan Guo, Jingcai Shao, Ling Tang, Hao Machine Learning Artificial Intelligence Computation and Language Large language models (LLMs) show impressive performance in solving complex language tasks. However, its large number of parameters presents significant challenges for the deployment. So, compressing LLMs to low bits can enable to deploy on resource-constrained devices. To address this problem, we propose gradient-aware weight quantization (GWQ), the first quantization approach for low-bit weight quantization that leverages gradients to localize outliers, requiring only a minimal amount of calibration data for outlier detection. GWQ retains the top 1\% outliers preferentially at FP16 precision, while the remaining non-outlier weights are stored in a low-bit. We widely evaluate GWQ on different task include language modeling, grounding detection, massive multitask language understanding and vision-language question and answering. Results show that models quantified by GWQ performs better than other quantization method. During quantization process, GWQ only need one calibration set to realize effective quant. Also, GWQ achieves 1.2x inference speedup in comparison to the original model and effectively reduces the inference memory. |
| title | GWQ: Gradient-Aware Weight Quantization for Large Language Models |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2411.00850 |