Post Training Quantization of Large Language Models with Microscaling Formats

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharify, Sayeh, Saxena, Utkarsh, Xu, Zifei, Yazar, Wanzin, Soloveychik, Ilya, Wang, Xin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909350816120832
author Sharify, Sayeh
Saxena, Utkarsh
Xu, Zifei
Yazar, Wanzin
Soloveychik, Ilya
Wang, Xin
author_facet Sharify, Sayeh
Saxena, Utkarsh
Xu, Zifei
Yazar, Wanzin
Soloveychik, Ilya
Wang, Xin
contents Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage challenges. This paper explores the potential of quantization to mitigate these challenges. We systematically study the combined application of three well-known post-training techniques, SmoothQuant, AWQ, and GPTQ, and provide a comprehensive analysis of their interactions and implications for advancing LLM quantization. We enhance the versatility of these methods by enabling quantization to microscaling (MX) formats, extending the applicability of these PTQ algorithms beyond their original fixed-point format targets. We show that combining different PTQ methods enables us to quantize models to 4-bit weights and 8-bit activations using the MXINT format with negligible accuracy loss compared to the uncompressed baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07135
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Post Training Quantization of Large Language Models with Microscaling Formats
Sharify, Sayeh
Saxena, Utkarsh
Xu, Zifei
Yazar, Wanzin
Soloveychik, Ilya
Wang, Xin
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage challenges. This paper explores the potential of quantization to mitigate these challenges. We systematically study the combined application of three well-known post-training techniques, SmoothQuant, AWQ, and GPTQ, and provide a comprehensive analysis of their interactions and implications for advancing LLM quantization. We enhance the versatility of these methods by enabling quantization to microscaling (MX) formats, extending the applicability of these PTQ algorithms beyond their original fixed-point format targets. We show that combining different PTQ methods enables us to quantize models to 4-bit weights and 8-bit activations using the MXINT format with negligible accuracy loss compared to the uncompressed baseline.
title Post Training Quantization of Large Language Models with Microscaling Formats
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.07135