BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Hao, Li, Lujun, Wang, Hao, Wang, Lei, Wang, Zheyu, Liu, Bei, Liu, Jiacheng, Zhu, Qiyuan, Han, Sirui, Guo, Yike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917393163354112
author Gu, Hao
Li, Lujun
Wang, Hao
Wang, Lei
Wang, Zheyu
Liu, Bei
Liu, Jiacheng
Zhu, Qiyuan
Han, Sirui
Guo, Yike
author_facet Gu, Hao
Li, Lujun
Wang, Hao
Wang, Lei
Wang, Zheyu
Liu, Bei
Liu, Jiacheng
Zhu, Qiyuan
Han, Sirui
Guo, Yike
contents Binary quantization represents the most extreme form of compression, reducing weights to +/-1 for maximal memory and computational efficiency. While recent sparsity-aware binarization achieves sub-1-bit compression via weight pruning, it faces critical challenges: performance degradation, mask-management overhead, and limited hardware compatibility. In this paper, we present BTC-LLM, a novel sub-1-bit LLM quantization framework that leverages binary pattern clustering and weight transformation to overcome these limitations. Our approach incorporates two key innovations: (1) a Binary Codebook that clusters recurring vectors into compact indices using custom distance metrics and sign-based updates; (2) a Learnable Transformation that reduces outliers and promotes shared sign patterns among binary weights. This eliminates sparse masks, enabling efficient inference on standard hardware. Extensive evaluations across LLaMA, Qwen, and FBI-LLM families demonstrate that BTC-LLM achieves state-of-the-art results in extreme compression (1.11-0.7 bits). Notably, BTC-LLM compressed to 0.8 bits on LLaMA-2-13B maintains high performance, with only a 3.1 percent accuracy drop in zero-shot benchmarks, while delivering a 1.6x speedup over FP16.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
Gu, Hao
Li, Lujun
Wang, Hao
Wang, Lei
Wang, Zheyu
Liu, Bei
Liu, Jiacheng
Zhu, Qiyuan
Han, Sirui
Guo, Yike
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Binary quantization represents the most extreme form of compression, reducing weights to +/-1 for maximal memory and computational efficiency. While recent sparsity-aware binarization achieves sub-1-bit compression via weight pruning, it faces critical challenges: performance degradation, mask-management overhead, and limited hardware compatibility. In this paper, we present BTC-LLM, a novel sub-1-bit LLM quantization framework that leverages binary pattern clustering and weight transformation to overcome these limitations. Our approach incorporates two key innovations: (1) a Binary Codebook that clusters recurring vectors into compact indices using custom distance metrics and sign-based updates; (2) a Learnable Transformation that reduces outliers and promotes shared sign patterns among binary weights. This eliminates sparse masks, enabling efficient inference on standard hardware. Extensive evaluations across LLaMA, Qwen, and FBI-LLM families demonstrate that BTC-LLM achieves state-of-the-art results in extreme compression (1.11-0.7 bits). Notably, BTC-LLM compressed to 0.8 bits on LLaMA-2-13B maintains high performance, with only a 3.1 percent accuracy drop in zero-shot benchmarks, while delivering a 1.6x speedup over FP16.
title BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.12040