DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Xiaoming, Tang, Shize, Yu, Guanghua, Xie, Linchuan, Liu, Song, Zhu, Jianchen, Li, Feng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912979091456000
author Yu, Xiaoming
Tang, Shize
Yu, Guanghua
Xie, Linchuan
Liu, Song
Zhu, Jianchen
Li, Feng
author_facet Yu, Xiaoming
Tang, Shize
Yu, Guanghua
Xie, Linchuan
Liu, Song
Zhu, Jianchen
Li, Feng
contents We introduce Delta-Aware Quantization (DAQ), a data-free post-training quantization framework that preserves the knowledge acquired during post-training. Standard quantization objectives minimize reconstruction error but are agnostic to the base model, allowing quantization noise to disproportionately corrupt the small-magnitude parameter deltas ($ΔW$) that encode post-training behavior -- an effect we analyze through the lens of quantization as implicit regularization. DAQ replaces reconstruction-based objectives with two delta-aware metrics -- Sign Preservation Rate and Cosine Similarity -- that directly optimize for directional fidelity of $ΔW$, requiring only the base and post-trained weight matrices. In a pilot FP8 study, DAQ recovers style-specific capabilities lost under standard quantization while maintaining general performance.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22324
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
Yu, Xiaoming
Tang, Shize
Yu, Guanghua
Xie, Linchuan
Liu, Song
Zhu, Jianchen
Li, Feng
Machine Learning
Artificial Intelligence
We introduce Delta-Aware Quantization (DAQ), a data-free post-training quantization framework that preserves the knowledge acquired during post-training. Standard quantization objectives minimize reconstruction error but are agnostic to the base model, allowing quantization noise to disproportionately corrupt the small-magnitude parameter deltas ($ΔW$) that encode post-training behavior -- an effect we analyze through the lens of quantization as implicit regularization. DAQ replaces reconstruction-based objectives with two delta-aware metrics -- Sign Preservation Rate and Cosine Similarity -- that directly optimize for directional fidelity of $ΔW$, requiring only the base and post-trained weight matrices. In a pilot FP8 study, DAQ recovers style-specific capabilities lost under standard quantization while maintaining general performance.
title DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.22324