FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Fangxin, Wang, Zongwu, Xia, JinHong, Zhao, Junping, Zhao, Shouren, Li, Jinjin, Liu, Jian, Jiang, Li, Guan, Haibing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912661818572800
author Liu, Fangxin
Wang, Zongwu
Xia, JinHong
Zhao, Junping
Zhao, Shouren
Li, Jinjin
Liu, Jian
Jiang, Li
Guan, Haibing
author_facet Liu, Fangxin
Wang, Zongwu
Xia, JinHong
Zhao, Junping
Zhao, Shouren
Li, Jinjin
Liu, Jian
Jiang, Li
Guan, Haibing
contents The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static quantization strategies, which struggle to adapt to dynamic workloads. To address this, we propose FlexQuant, a dynamic precision-switching framework that optimizes the trade-off between inference speed and accuracy. Leveraging model perplexity entropy and Kullback-Leibler divergence, FlexQuant enables fine-grained, layer-wise mixed-precision quantization and dynamically adjusts bit-widths during each token generation. FlexQuant provides a comprehensive analysis of quantization strategies, introduces a precision requirement model for optimal switching, and implements efficient fine-grained precision management. Evaluations demonstrate that FlexQuant achieves a 1.3x end-to-end speedup across diverse language tasks with negligible accuracy loss introduced. This framework offers a flexible and adaptive solution for efficient LLM deployment. Code is released at https://github.com/ZongwuWang/FlexQuant.git.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
Liu, Fangxin
Wang, Zongwu
Xia, JinHong
Zhao, Junping
Zhao, Shouren
Li, Jinjin
Liu, Jian
Jiang, Li
Guan, Haibing
Machine Learning
I.2.1; I.2.7
The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static quantization strategies, which struggle to adapt to dynamic workloads. To address this, we propose FlexQuant, a dynamic precision-switching framework that optimizes the trade-off between inference speed and accuracy. Leveraging model perplexity entropy and Kullback-Leibler divergence, FlexQuant enables fine-grained, layer-wise mixed-precision quantization and dynamically adjusts bit-widths during each token generation. FlexQuant provides a comprehensive analysis of quantization strategies, introduces a precision requirement model for optimal switching, and implements efficient fine-grained precision management. Evaluations demonstrate that FlexQuant achieves a 1.3x end-to-end speedup across diverse language tasks with negligible accuracy loss introduced. This framework offers a flexible and adaptive solution for efficient LLM deployment. Code is released at https://github.com/ZongwuWang/FlexQuant.git.
title FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
topic Machine Learning
I.2.1; I.2.7
url https://arxiv.org/abs/2506.12024