D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Xianglong, Bao, ChengZhu, Li, Zhiteng, Zhang, Tianao, Zhang, Shaoqiu, Xie, Ruobing, Sun, Samm, Zhang, Yulun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914310304825344
author Yan, Xianglong
Bao, ChengZhu
Li, Zhiteng
Zhang, Tianao
Zhang, Shaoqiu
Xie, Ruobing
Sun, Samm
Zhang, Yulun
author_facet Yan, Xianglong
Bao, ChengZhu
Li, Zhiteng
Zhang, Tianao
Zhang, Shaoqiu
Xie, Ruobing
Sun, Samm
Zhang, Yulun
contents Large language models (LLMs) deliver strong performance, but their high compute and memory costs make deployment difficult in resource-constrained scenarios. Weight-only post-training quantization (PTQ) is appealing, as it reduces memory usage and enables practical speedup without low-bit operators or specialized hardware. However, accuracy often degrades significantly in weight-only PTQ at sub-4-bit precision, and our analysis identifies two main causes: (1) down-projection matrices are a well-known quantization bottleneck, but maintaining their fidelity often requires extra bit-width; (2) weight quantization induces activation deviations, but effective correction strategies remain underexplored. To address these issues, we propose D$^2$Quant, a novel weight-only PTQ framework that improves quantization from both the weight and activation perspectives. On the weight side, we design a Dual-Scale Quantizer (DSQ) tailored to down-projection matrices, with an absorbable scaling factor that significantly improves accuracy without increasing the bit budget. On the activation side, we propose Deviation-Aware Correction (DAC), which incorporates a mean-shift correction within LayerNorm to mitigate quantization-induced activation distribution shifts. Extensive experiments across multiple LLM families and evaluation metrics show that D$^2$Quant delivers superior performance for weight-only PTQ at sub-4-bit precision. The code and models will be available at https://github.com/XIANGLONGYAN/D2Quant.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02546
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
Yan, Xianglong
Bao, ChengZhu
Li, Zhiteng
Zhang, Tianao
Zhang, Shaoqiu
Xie, Ruobing
Sun, Samm
Zhang, Yulun
Machine Learning
Artificial Intelligence
Large language models (LLMs) deliver strong performance, but their high compute and memory costs make deployment difficult in resource-constrained scenarios. Weight-only post-training quantization (PTQ) is appealing, as it reduces memory usage and enables practical speedup without low-bit operators or specialized hardware. However, accuracy often degrades significantly in weight-only PTQ at sub-4-bit precision, and our analysis identifies two main causes: (1) down-projection matrices are a well-known quantization bottleneck, but maintaining their fidelity often requires extra bit-width; (2) weight quantization induces activation deviations, but effective correction strategies remain underexplored. To address these issues, we propose D$^2$Quant, a novel weight-only PTQ framework that improves quantization from both the weight and activation perspectives. On the weight side, we design a Dual-Scale Quantizer (DSQ) tailored to down-projection matrices, with an absorbable scaling factor that significantly improves accuracy without increasing the bit budget. On the activation side, we propose Deviation-Aware Correction (DAC), which incorporates a mean-shift correction within LayerNorm to mitigate quantization-induced activation distribution shifts. Extensive experiments across multiple LLM families and evaluation metrics show that D$^2$Quant delivers superior performance for weight-only PTQ at sub-4-bit precision. The code and models will be available at https://github.com/XIANGLONGYAN/D2Quant.
title D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.02546