DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912979091456000 |
|---|---|
| author | Yu, Xiaoming Tang, Shize Yu, Guanghua Xie, Linchuan Liu, Song Zhu, Jianchen Li, Feng |
| author_facet | Yu, Xiaoming Tang, Shize Yu, Guanghua Xie, Linchuan Liu, Song Zhu, Jianchen Li, Feng |
| contents | We introduce Delta-Aware Quantization (DAQ), a data-free post-training quantization framework that preserves the knowledge acquired during post-training. Standard quantization objectives minimize reconstruction error but are agnostic to the base model, allowing quantization noise to disproportionately corrupt the small-magnitude parameter deltas ($ΔW$) that encode post-training behavior -- an effect we analyze through the lens of quantization as implicit regularization. DAQ replaces reconstruction-based objectives with two delta-aware metrics -- Sign Preservation Rate and Cosine Similarity -- that directly optimize for directional fidelity of $ΔW$, requiring only the base and post-trained weight matrices. In a pilot FP8 study, DAQ recovers style-specific capabilities lost under standard quantization while maintaining general performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_22324 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression Yu, Xiaoming Tang, Shize Yu, Guanghua Xie, Linchuan Liu, Song Zhu, Jianchen Li, Feng Machine Learning Artificial Intelligence We introduce Delta-Aware Quantization (DAQ), a data-free post-training quantization framework that preserves the knowledge acquired during post-training. Standard quantization objectives minimize reconstruction error but are agnostic to the base model, allowing quantization noise to disproportionately corrupt the small-magnitude parameter deltas ($ΔW$) that encode post-training behavior -- an effect we analyze through the lens of quantization as implicit regularization. DAQ replaces reconstruction-based objectives with two delta-aware metrics -- Sign Preservation Rate and Cosine Similarity -- that directly optimize for directional fidelity of $ΔW$, requiring only the base and post-trained weight matrices. In a pilot FP8 study, DAQ recovers style-specific capabilities lost under standard quantization while maintaining general performance. |
| title | DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2603.22324 |