Saliency-Aware Regularized Quantization Calibration for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yanlong, Cheng, Xiaoyuan, Liu, Huihang, He, Baihua, Zhang, Xinyu, Zhu, Harrison Bo Hua, Chen, Wenlong, Zeng, Li, Sun, Zhuo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909025225932800
author Zhao, Yanlong
Cheng, Xiaoyuan
Liu, Huihang
He, Baihua
Zhang, Xinyu
Zhu, Harrison Bo Hua
Chen, Wenlong
Zeng, Li
Sun, Zhuo
author_facet Zhao, Yanlong
Cheng, Xiaoyuan
Liu, Huihang
He, Baihua
Zhang, Xinyu
Zhu, Harrison Bo Hua
Chen, Wenlong
Zeng, Li
Sun, Zhuo
contents Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most existing PTQ methods determine quantization parameters by minimizing a layer-wise reconstruction error on a predetermined calibration dataset, typically optimized via either scale search or Gram-based methods. However, from the perspective of generalization risk, existing PTQ calibration objectives based solely on empirical reconstruction error over limited or unrepresentative calibration data may move the quantized weights away from the original floating-point weights, potentially degrading downstream performance. To address this issue, we propose \emph{Regularized Quantization Calibration} (RQC), a unified framework that augments standard PTQ objectives with a regularizer that explicitly controls weight deviation from the original weights. We further generalize this framework to incorporate a saliency-aware regularizer, resulting in \emph{Saliency-Aware Regularized Quantization Calibration} (SARQC). The proposed regularization encourages quantized weights to remain close to the original weights during calibration, leading to improved generalization at inference time. SARQC integrates seamlessly into existing PTQ pipelines and enhances both scale-search-based and Gram-based methods under a unified formulation. Extensive experiments on dense and Mixture-of-Experts LLMs demonstrate consistent improvements in perplexity and zero-shot accuracy, without introducing additional inference overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05693
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Saliency-Aware Regularized Quantization Calibration for Large Language Models
Zhao, Yanlong
Cheng, Xiaoyuan
Liu, Huihang
He, Baihua
Zhang, Xinyu
Zhu, Harrison Bo Hua
Chen, Wenlong
Zeng, Li
Sun, Zhuo
Artificial Intelligence
Machine Learning
Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most existing PTQ methods determine quantization parameters by minimizing a layer-wise reconstruction error on a predetermined calibration dataset, typically optimized via either scale search or Gram-based methods. However, from the perspective of generalization risk, existing PTQ calibration objectives based solely on empirical reconstruction error over limited or unrepresentative calibration data may move the quantized weights away from the original floating-point weights, potentially degrading downstream performance. To address this issue, we propose \emph{Regularized Quantization Calibration} (RQC), a unified framework that augments standard PTQ objectives with a regularizer that explicitly controls weight deviation from the original weights. We further generalize this framework to incorporate a saliency-aware regularizer, resulting in \emph{Saliency-Aware Regularized Quantization Calibration} (SARQC). The proposed regularization encourages quantized weights to remain close to the original weights during calibration, leading to improved generalization at inference time. SARQC integrates seamlessly into existing PTQ pipelines and enhances both scale-search-based and Gram-based methods under a unified formulation. Extensive experiments on dense and Mixture-of-Experts LLMs demonstrate consistent improvements in perplexity and zero-shot accuracy, without introducing additional inference overhead.
title Saliency-Aware Regularized Quantization Calibration for Large Language Models
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.05693