Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Jaemin, Kim, Sungkyun, Lee, Junyeol, Seo, Jiwon
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910131413843968
author Kim, Jaemin
Kim, Sungkyun
Lee, Junyeol
Seo, Jiwon
author_facet Kim, Jaemin
Kim, Sungkyun
Lee, Junyeol
Seo, Jiwon
contents Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
Kim, Jaemin
Kim, Sungkyun
Lee, Junyeol
Seo, Jiwon
Machine Learning
Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.
title Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
topic Machine Learning
url https://arxiv.org/abs/2604.13806