Post Training Quantization of Large Language Models with Microscaling Formats
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909350816120832 |
|---|---|
| author | Sharify, Sayeh Saxena, Utkarsh Xu, Zifei Yazar, Wanzin Soloveychik, Ilya Wang, Xin |
| author_facet | Sharify, Sayeh Saxena, Utkarsh Xu, Zifei Yazar, Wanzin Soloveychik, Ilya Wang, Xin |
| contents | Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage challenges. This paper explores the potential of quantization to mitigate these challenges. We systematically study the combined application of three well-known post-training techniques, SmoothQuant, AWQ, and GPTQ, and provide a comprehensive analysis of their interactions and implications for advancing LLM quantization. We enhance the versatility of these methods by enabling quantization to microscaling (MX) formats, extending the applicability of these PTQ algorithms beyond their original fixed-point format targets. We show that combining different PTQ methods enables us to quantize models to 4-bit weights and 8-bit activations using the MXINT format with negligible accuracy loss compared to the uncompressed baseline. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_07135 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Post Training Quantization of Large Language Models with Microscaling Formats Sharify, Sayeh Saxena, Utkarsh Xu, Zifei Yazar, Wanzin Soloveychik, Ilya Wang, Xin Machine Learning Artificial Intelligence Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage challenges. This paper explores the potential of quantization to mitigate these challenges. We systematically study the combined application of three well-known post-training techniques, SmoothQuant, AWQ, and GPTQ, and provide a comprehensive analysis of their interactions and implications for advancing LLM quantization. We enhance the versatility of these methods by enabling quantization to microscaling (MX) formats, extending the applicability of these PTQ algorithms beyond their original fixed-point format targets. We show that combining different PTQ methods enables us to quantize models to 4-bit weights and 8-bit activations using the MXINT format with negligible accuracy loss compared to the uncompressed baseline. |
| title | Post Training Quantization of Large Language Models with Microscaling Formats |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2405.07135 |