Leech Lattice Vector Quantization for Efficient LLM Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: van der Ouderaa, Tycho F. A., van Baalen, Mart, Whatmough, Paul, Nagel, Markus
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918383521366016
author van der Ouderaa, Tycho F. A.
van Baalen, Mart
Whatmough, Paul
Nagel, Markus
author_facet van der Ouderaa, Tycho F. A.
van Baalen, Mart
Whatmough, Paul
Nagel, Markus
contents Scalar quantization of large language models (LLMs) is fundamentally limited by information-theoretic bounds. While vector quantization (VQ) overcomes these limits by encoding blocks of parameters jointly, practical implementations must avoid the need for expensive lookup mechanisms or other explicit codebook storage. Lattice approaches address this through highly structured and dense packing. This paper explores the Leech lattice, which, with its optimal sphere packing and kissing configurations at 24 dimensions, is the highest dimensional lattice known with such optimal properties. To make the Leech lattice usable for LLM quantization, we extend an existing search algorithm based on the extended Golay code construction, to i) support indexing, enabling conversion to and from bitstrings without materializing the codebook, ii) allow angular search over union of Leech lattice shells, iii) propose fully-parallelisable dequantization kernel. Together this yields a practical algorithm, namely Leech Lattice Vector Quantization (LLVQ). LLVQ delivers state-of-the-art LLM quantization performance, outperforming recent methods such as Quip\#, QTIP, and PVQ. These results highlight the importance of high-dimensional lattices for scalable, theoretically grounded model compression.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11021
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Leech Lattice Vector Quantization for Efficient LLM Compression
van der Ouderaa, Tycho F. A.
van Baalen, Mart
Whatmough, Paul
Nagel, Markus
Machine Learning
Scalar quantization of large language models (LLMs) is fundamentally limited by information-theoretic bounds. While vector quantization (VQ) overcomes these limits by encoding blocks of parameters jointly, practical implementations must avoid the need for expensive lookup mechanisms or other explicit codebook storage. Lattice approaches address this through highly structured and dense packing. This paper explores the Leech lattice, which, with its optimal sphere packing and kissing configurations at 24 dimensions, is the highest dimensional lattice known with such optimal properties. To make the Leech lattice usable for LLM quantization, we extend an existing search algorithm based on the extended Golay code construction, to i) support indexing, enabling conversion to and from bitstrings without materializing the codebook, ii) allow angular search over union of Leech lattice shells, iii) propose fully-parallelisable dequantization kernel. Together this yields a practical algorithm, namely Leech Lattice Vector Quantization (LLVQ). LLVQ delivers state-of-the-art LLM quantization performance, outperforming recent methods such as Quip\#, QTIP, and PVQ. These results highlight the importance of high-dimensional lattices for scalable, theoretically grounded model compression.
title Leech Lattice Vector Quantization for Efficient LLM Compression
topic Machine Learning
url https://arxiv.org/abs/2603.11021