Continuous Approximations for Improving Quantization Aware Training of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, He, Hong, Jianhang, Wu, Yuanzhuo, Adbol, Snehal, Li, Zonglin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913545496559616
author Li, He
Hong, Jianhang
Wu, Yuanzhuo
Adbol, Snehal
Li, Zonglin
author_facet Li, He
Hong, Jianhang
Wu, Yuanzhuo
Adbol, Snehal
Li, Zonglin
contents Model compression methods are used to reduce the computation and energy requirements for Large Language Models (LLMs). Quantization Aware Training (QAT), an effective model compression method, is proposed to reduce performance degradation after quantization. To further minimize this degradation, we introduce two continuous approximations to the QAT process on the rounding function, traditionally approximated by the Straight-Through Estimator (STE), and the clamping function. By applying both methods, the perplexity (PPL) on the WikiText-v2 dataset of the quantized model reaches 9.0815, outperforming 9.9621 by the baseline. Also, we achieve a 2.76% improvement on BoolQ, and a 5.47% improvement on MMLU, proving that the step sizes and weights can be learned more accurately with our approach. Our method achieves better performance with the same precision, model size, and training setup, contributing to the development of more energy-efficient LLMs technology that aligns with global sustainability goals.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10849
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous Approximations for Improving Quantization Aware Training of LLMs
Li, He
Hong, Jianhang
Wu, Yuanzhuo
Adbol, Snehal
Li, Zonglin
Machine Learning
Artificial Intelligence
Computation and Language
Model compression methods are used to reduce the computation and energy requirements for Large Language Models (LLMs). Quantization Aware Training (QAT), an effective model compression method, is proposed to reduce performance degradation after quantization. To further minimize this degradation, we introduce two continuous approximations to the QAT process on the rounding function, traditionally approximated by the Straight-Through Estimator (STE), and the clamping function. By applying both methods, the perplexity (PPL) on the WikiText-v2 dataset of the quantized model reaches 9.0815, outperforming 9.9621 by the baseline. Also, we achieve a 2.76% improvement on BoolQ, and a 5.47% improvement on MMLU, proving that the step sizes and weights can be learned more accurately with our approach. Our method achieves better performance with the same precision, model size, and training setup, contributing to the development of more energy-efficient LLMs technology that aligns with global sustainability goals.
title Continuous Approximations for Improving Quantization Aware Training of LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.10849