Extreme Compression of Large Language Models via Additive Quantization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Egiazarian, Vage, Panferov, Andrei, Kuznedelev, Denis, Frantar, Elias, Babenko, Artem, Alistarh, Dan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914946960326656
author Egiazarian, Vage
Panferov, Andrei
Kuznedelev, Denis
Frantar, Elias
Babenko, Artem
Alistarh, Dan
author_facet Egiazarian, Vage
Panferov, Andrei
Kuznedelev, Denis
Frantar, Elias
Babenko, Artem
Alistarh, Dan
contents The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of "extreme" LLM compression-defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter-from the point of view of classic methods in Multi-Codebook Quantization (MCQ). Our algorithm, called AQLM, generalizes the classic Additive Quantization (AQ) approach for information retrieval to advance the state-of-the-art in LLM compression, via two innovations: 1) learned additive quantization of weight matrices in input-adaptive fashion, and 2) joint optimization of codebook parameters across each transformer blocks. Broadly, AQLM is the first scheme that is Pareto optimal in terms of accuracy-vs-model-size when compressing to less than 3 bits per parameter, and significantly improves upon all known schemes in the extreme compression (2bit) regime. In addition, AQLM is practical: we provide fast GPU and CPU implementations of AQLM for token generation, which enable us to match or outperform optimized FP16 implementations for speed, while executing in a much smaller memory footprint.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06118
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extreme Compression of Large Language Models via Additive Quantization
Egiazarian, Vage
Panferov, Andrei
Kuznedelev, Denis
Frantar, Elias
Babenko, Artem
Alistarh, Dan
Machine Learning
Computation and Language
The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of "extreme" LLM compression-defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter-from the point of view of classic methods in Multi-Codebook Quantization (MCQ). Our algorithm, called AQLM, generalizes the classic Additive Quantization (AQ) approach for information retrieval to advance the state-of-the-art in LLM compression, via two innovations: 1) learned additive quantization of weight matrices in input-adaptive fashion, and 2) joint optimization of codebook parameters across each transformer blocks. Broadly, AQLM is the first scheme that is Pareto optimal in terms of accuracy-vs-model-size when compressing to less than 3 bits per parameter, and significantly improves upon all known schemes in the extreme compression (2bit) regime. In addition, AQLM is practical: we provide fast GPU and CPU implementations of AQLM for token generation, which enable us to match or outperform optimized FP16 implementations for speed, while executing in a much smaller memory footprint.
title Extreme Compression of Large Language Models via Additive Quantization
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2401.06118