LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Fangxin, Yang, Ning, Zhao, Junping, Yang, Tao, Guan, Haibing, Jiang, Li
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916793807798272
author Liu, Fangxin
Yang, Ning
Zhao, Junping
Yang, Tao
Guan, Haibing
Jiang, Li
author_facet Liu, Fangxin
Yang, Ning
Zhao, Junping
Yang, Tao
Guan, Haibing
Jiang, Li
contents Large language models (LLMs) have achieved significant progress in natural language processing but face challenges in deployment due to high memory and computational requirements. Weight quantization is a common approach to address these issues, yet achieving effective low-bit compression remains challenging. This paper presents LCD, which unifies the learning of clustering-based quantization within a knowledge distillation framework. Using carefully designed optimization techniques, LCD preserves LLM performance even at ultra-low bit widths of 2-3 bits. Additionally, LCD compresses activations through smoothing and accelerates inference with a LUT-based design. Experimental results show that LCD outperforms existing methods and delivers up to a 6.2x speedup in inference. Notably, LCD is shown to be more cost-effective, making it a practical solution for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
Liu, Fangxin
Yang, Ning
Zhao, Junping
Yang, Tao
Guan, Haibing
Jiang, Li
Machine Learning
Artificial Intelligence
Large language models (LLMs) have achieved significant progress in natural language processing but face challenges in deployment due to high memory and computational requirements. Weight quantization is a common approach to address these issues, yet achieving effective low-bit compression remains challenging. This paper presents LCD, which unifies the learning of clustering-based quantization within a knowledge distillation framework. Using carefully designed optimization techniques, LCD preserves LLM performance even at ultra-low bit widths of 2-3 bits. Additionally, LCD compresses activations through smoothing and accelerates inference with a LUT-based design. Experimental results show that LCD outperforms existing methods and delivers up to a 6.2x speedup in inference. Notably, LCD is shown to be more cost-effective, making it a practical solution for real-world applications.
title LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.12038