JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yi, Kai, Vivekraja, Vignesh, Khaitan, Harshit, Li, Steven
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916044606537728
author Yi, Kai
Vivekraja, Vignesh
Khaitan, Harshit
Li, Steven
author_facet Yi, Kai
Vivekraja, Vignesh
Khaitan, Harshit
Li, Steven
contents Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-differentiable quantizers by fiat. This often makes training brittle near bin boundaries and weakly aligned with the actual behavior of the low-precision model. We introduce JacQuant, a QAT framework that learns a lightweight surrogate of the model's local sensitivity to parameter changes and uses it to stabilize and accelerate training within standard variance-reduced optimizers. The surrogate is inexpensive (diagonal or block-diagonal), data-driven, and compatible with common weight and activation quantizers. On code-preserving training phases, we prove convergence for non-convex objectives and obtain linear rates under a PL condition, and we relate the learned sensitivity to end-to-end output fidelity via a simple calibration argument. Across LLM benchmarks at $\leq 2$ bits, JacQuant consistently reaches higher accuracy than STE-based QAT, and the runtime analyses on various models show that the added cost remains negligible under practical group sizes. The method is drop-in and requires no changes to the forward quantizers; our empirical claims are scoped to ultra-low-bit LLM QAT.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25469
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates
Yi, Kai
Vivekraja, Vignesh
Khaitan, Harshit
Li, Steven
Machine Learning
Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-differentiable quantizers by fiat. This often makes training brittle near bin boundaries and weakly aligned with the actual behavior of the low-precision model. We introduce JacQuant, a QAT framework that learns a lightweight surrogate of the model's local sensitivity to parameter changes and uses it to stabilize and accelerate training within standard variance-reduced optimizers. The surrogate is inexpensive (diagonal or block-diagonal), data-driven, and compatible with common weight and activation quantizers. On code-preserving training phases, we prove convergence for non-convex objectives and obtain linear rates under a PL condition, and we relate the learned sensitivity to end-to-end output fidelity via a simple calibration argument. Across LLM benchmarks at $\leq 2$ bits, JacQuant consistently reaches higher accuracy than STE-based QAT, and the runtime analyses on various models show that the added cost remains negligible under practical group sizes. The method is drop-in and requires no changes to the forward quantizers; our empirical claims are scoped to ultra-low-bit LLM QAT.
title JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates
topic Machine Learning
url https://arxiv.org/abs/2605.25469