Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Hasan, Jahid
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915012471160832
author Hasan, Jahid
author_facet Hasan, Jahid
contents This paper presents a comprehensive analysis of quantization techniques for optimizing Large Language Models (LLMs), specifically focusing on Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT). Through empirical evaluation across models ranging from 10M to 1B parameters, we demonstrate that quantization can achieve up to 68% reduction in model size while maintaining performance within 6% of full-precision baselines when utilizing our proposed scaling factor γ. Our experiments show that INT8 quantization delivers a 40% reduction in computational cost and power consumption, while INT4 quantization further improves these metrics by 60%. We introduce a novel theoretical framework for mixed-precision quantization, deriving optimal bit allocation strategies based on layer sensitivity and weight variance. Hardware efficiency evaluations on edge devices reveal that our quantization approach enables up to 2.4x throughput improvement for INT8 and 3x for INT4, with 60% power reduction compared to full-precision models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06084
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
Hasan, Jahid
Machine Learning
Artificial Intelligence
Computation and Language
This paper presents a comprehensive analysis of quantization techniques for optimizing Large Language Models (LLMs), specifically focusing on Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT). Through empirical evaluation across models ranging from 10M to 1B parameters, we demonstrate that quantization can achieve up to 68% reduction in model size while maintaining performance within 6% of full-precision baselines when utilizing our proposed scaling factor γ. Our experiments show that INT8 quantization delivers a 40% reduction in computational cost and power consumption, while INT4 quantization further improves these metrics by 60%. We introduce a novel theoretical framework for mixed-precision quantization, deriving optimal bit allocation strategies based on layer sensitivity and weight variance. Hardware efficiency evaluations on edge devices reveal that our quantization approach enables up to 2.4x throughput improvement for INT8 and 3x for INT4, with 60% power reduction compared to full-precision models.
title Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.06084