Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Seungcheol, Bae, Jeongin, Kwon, Beomseok, Kim, Minjun, Kim, Byeongwook, Kwon, Se Jung, Kang, U, Lee, Dongsoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912431493611520
author Park, Seungcheol
Bae, Jeongin
Kwon, Beomseok
Kim, Minjun
Kim, Byeongwook
Kwon, Se Jung
Kang, U
Lee, Dongsoo
author_facet Park, Seungcheol
Bae, Jeongin
Kwon, Beomseok
Kim, Minjun
Kim, Byeongwook
Kwon, Se Jung
Kang, U
Lee, Dongsoo
contents How can we quantize large language models while preserving accuracy? Quantization is essential for deploying large language models (LLMs) efficiently. Binary-coding quantization (BCQ) and uniform quantization (UQ) are promising quantization schemes that have strong expressiveness and optimizability, respectively. However, neither scheme leverages both advantages. In this paper, we propose UniQuanF (Unified Quantization with Flexible Mapping), an accurate quantization method for LLMs. UniQuanF harnesses both strong expressiveness and optimizability by unifying the flexible mapping technique in UQ and non-uniform quantization levels of BCQ. We propose unified initialization, and local and periodic mapping techniques to optimize the parameters in UniQuanF precisely. After optimization, our unification theorem removes computational and memory overhead, allowing us to utilize the superior accuracy of UniQuanF without extra deployment costs induced by the unification. Experimental results demonstrate that UniQuanF outperforms existing UQ and BCQ methods, achieving up to 4.60% higher accuracy on GSM8K benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
Park, Seungcheol
Bae, Jeongin
Kwon, Beomseok
Kim, Minjun
Kim, Byeongwook
Kwon, Se Jung
Kang, U
Lee, Dongsoo
Computation and Language
68T50
I.2.7
How can we quantize large language models while preserving accuracy? Quantization is essential for deploying large language models (LLMs) efficiently. Binary-coding quantization (BCQ) and uniform quantization (UQ) are promising quantization schemes that have strong expressiveness and optimizability, respectively. However, neither scheme leverages both advantages. In this paper, we propose UniQuanF (Unified Quantization with Flexible Mapping), an accurate quantization method for LLMs. UniQuanF harnesses both strong expressiveness and optimizability by unifying the flexible mapping technique in UQ and non-uniform quantization levels of BCQ. We propose unified initialization, and local and periodic mapping techniques to optimize the parameters in UniQuanF precisely. After optimization, our unification theorem removes computational and memory overhead, allowing us to utilize the superior accuracy of UniQuanF without extra deployment costs induced by the unification. Experimental results demonstrate that UniQuanF outperforms existing UQ and BCQ methods, achieving up to 4.60% higher accuracy on GSM8K benchmark.
title Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2506.03781