InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ke, An, Dong, Zang, Xiaoling, Ye, Can, Xie, Liang, Qiu, Qibo, Shen, Chen, He, Xiaofei, Wang, Wenxiao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910256962994176
author Li, Ke
An, Dong
Zang, Xiaoling
Ye, Can
Xie, Liang
Qiu, Qibo
Shen, Chen
He, Xiaofei
Wang, Wenxiao
author_facet Li, Ke
An, Dong
Zang, Xiaoling
Ye, Can
Xie, Liang
Qiu, Qibo
Shen, Chen
He, Xiaofei
Wang, Wenxiao
contents Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activations contain outliers, but that their distributions are often poorly matched to a low-bit uniform quantizer. Existing post-training quantization (PTQ) methods suppress peaks, balance channels, or minimize reconstruction error, yet they rarely specify what activation distribution is actually easy to discretize. As a result, activations may appear numerically smoother while still incurring large quantization error because the quantization range remains wide or most values collapse into a few levels near the mean. We recast activation transformation as quantizer-facing distribution design and analyze quantization error from an information-theoretic perspective. Our analysis shows that quantization-friendly activations should jointly have a smaller numerical range and sufficient dispersion within that range. Guided by this analysis, we propose InfoQuant, a train-free method that employs Peak Suppression Orthogonal Transformation (PSOT) to shape activations into more quantization-friendly distributions. We further introduce adaptive outlier-token selection to improve the robustness of PSOT during optimization. Across multiple LLM families, InfoQuant consistently outperforms prior PTQ and end-to-end training baselines. Under W4A4KV4, it preserves 97% of floating-point accuracy on average and reduces the LLaMA-2 13B performance gap by 42% over the previous state of the art. Code is available at [https://github.com/LLIKKE/InfoQuant](https://github.com/LLIKKE/InfoQuant)
format Preprint
id arxiv_https___arxiv_org_abs_2605_26175
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
Li, Ke
An, Dong
Zang, Xiaoling
Ye, Can
Xie, Liang
Qiu, Qibo
Shen, Chen
He, Xiaofei
Wang, Wenxiao
Machine Learning
Artificial Intelligence
Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activations contain outliers, but that their distributions are often poorly matched to a low-bit uniform quantizer. Existing post-training quantization (PTQ) methods suppress peaks, balance channels, or minimize reconstruction error, yet they rarely specify what activation distribution is actually easy to discretize. As a result, activations may appear numerically smoother while still incurring large quantization error because the quantization range remains wide or most values collapse into a few levels near the mean. We recast activation transformation as quantizer-facing distribution design and analyze quantization error from an information-theoretic perspective. Our analysis shows that quantization-friendly activations should jointly have a smaller numerical range and sufficient dispersion within that range. Guided by this analysis, we propose InfoQuant, a train-free method that employs Peak Suppression Orthogonal Transformation (PSOT) to shape activations into more quantization-friendly distributions. We further introduce adaptive outlier-token selection to improve the robustness of PSOT during optimization. Across multiple LLM families, InfoQuant consistently outperforms prior PTQ and end-to-end training baselines. Under W4A4KV4, it preserves 97% of floating-point accuracy on average and reduces the LLaMA-2 13B performance gap by 42% over the previous state of the art. Code is available at [https://github.com/LLIKKE/InfoQuant](https://github.com/LLIKKE/InfoQuant)
title InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.26175