QuIP: 2-Bit Quantization of Large Language Models With Guarantees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chee, Jerry, Cai, Yaohui, Kuleshov, Volodymyr, De Sa, Christopher
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929209336659968
author Chee, Jerry
Cai, Yaohui
Kuleshov, Volodymyr
De Sa, Christopher
author_facet Chee, Jerry
Cai, Yaohui
Kuleshov, Volodymyr
De Sa, Christopher
contents This work studies post-training parameter quantization in large language models (LLMs). We introduce quantization with incoherence processing (QuIP), a new method based on the insight that quantization benefits from $\textit{incoherent}$ weight and Hessian matrices, i.e., from the weights being even in magnitude and the directions in which it is important to round them accurately being unaligned with the coordinate axes. QuIP consists of two steps: (1) an adaptive rounding procedure minimizing a quadratic proxy objective; (2) efficient pre- and post-processing that ensures weight and Hessian incoherence via multiplication by random orthogonal matrices. We complement QuIP with the first theoretical analysis for an LLM-scale quantization algorithm, and show that our theory also applies to an existing method, OPTQ. Empirically, we find that our incoherence preprocessing improves several existing quantization algorithms and yields the first LLM quantization methods that produce viable results using only two bits per weight. Our code can be found at https://github.com/Cornell-RelaxML/QuIP.
format Preprint
id arxiv_https___arxiv_org_abs_2307_13304
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Chee, Jerry
Cai, Yaohui
Kuleshov, Volodymyr
De Sa, Christopher
Machine Learning
Computation and Language
This work studies post-training parameter quantization in large language models (LLMs). We introduce quantization with incoherence processing (QuIP), a new method based on the insight that quantization benefits from $\textit{incoherent}$ weight and Hessian matrices, i.e., from the weights being even in magnitude and the directions in which it is important to round them accurately being unaligned with the coordinate axes. QuIP consists of two steps: (1) an adaptive rounding procedure minimizing a quadratic proxy objective; (2) efficient pre- and post-processing that ensures weight and Hessian incoherence via multiplication by random orthogonal matrices. We complement QuIP with the first theoretical analysis for an LLM-scale quantization algorithm, and show that our theory also applies to an existing method, OPTQ. Empirically, we find that our incoherence preprocessing improves several existing quantization algorithms and yields the first LLM quantization methods that produce viable results using only two bits per weight. Our code can be found at https://github.com/Cornell-RelaxML/QuIP.
title QuIP: 2-Bit Quantization of Large Language Models With Guarantees
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2307.13304