Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Wenhua, Zhang, Weiwei, Shen, Haihao, Cai, Yiyang, He, Xin, Lv, Kaokao, Liu, Yi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912062283710464
author Cheng, Wenhua
Zhang, Weiwei
Shen, Haihao
Cai, Yiyang
He, Xin
Lv, Kaokao
Liu, Yi
author_facet Cheng, Wenhua
Zhang, Weiwei
Shen, Haihao
Cai, Yiyang
He, Xin
Lv, Kaokao
Liu, Yi
contents Large Language Models (LLMs) have demonstrated exceptional proficiency in language-related tasks, but their deployment poses significant challenges due to substantial memory and storage requirements. Weight-only quantization has emerged as a promising solution, significantly reducing memory and storage needs without sacrificing too much performance. In this study, we introduce SignRound, a method that leverages signed gradient descent (SignSGD) to optimize rounding values and weight clipping in just 200 steps. SignRound integrates the advantages of Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ), delivering exceptional results across 2 to 4 bits while minimizing tuning costs and avoiding additional inference overhead. For example, SignRound achieved absolute average accuracy improvements ranging from 6.91% to 33.22% at 2bits, as measured by the average zero-shot accuracy across 11 tasks. It also demonstrates strong generalization in recent models, achieving near-lossless 4-bit quantization in most scenarios. The source code is publicly available at https://github.com/intel/auto-round.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05516
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
Cheng, Wenhua
Zhang, Weiwei
Shen, Haihao
Cai, Yiyang
He, Xin
Lv, Kaokao
Liu, Yi
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) have demonstrated exceptional proficiency in language-related tasks, but their deployment poses significant challenges due to substantial memory and storage requirements. Weight-only quantization has emerged as a promising solution, significantly reducing memory and storage needs without sacrificing too much performance. In this study, we introduce SignRound, a method that leverages signed gradient descent (SignSGD) to optimize rounding values and weight clipping in just 200 steps. SignRound integrates the advantages of Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ), delivering exceptional results across 2 to 4 bits while minimizing tuning costs and avoiding additional inference overhead. For example, SignRound achieved absolute average accuracy improvements ranging from 6.91% to 33.22% at 2bits, as measured by the average zero-shot accuracy across 11 tasks. It also demonstrates strong generalization in recent models, achieving near-lossless 4-bit quantization in most scenarios. The source code is publicly available at https://github.com/intel/auto-round.
title Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2309.05516