DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Yingsong, Chen, Ling
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912074976722944
author Luo, Yingsong
Chen, Ling
author_facet Luo, Yingsong
Chen, Ling
contents Large language models (LLMs) excel in various tasks but face deployment challenges due to hardware constraints. We propose density-aware post-training weight-only quantization (DAQ), which has two stages: 1) density-centric alignment, which identifies the center of high-density weights and centers the dynamic range on this point to align high-density weight regions with floating-point high-precision regions; 2) learnable dynamic range adjustment, which adjusts the dynamic range by optimizing quantization parameters (i.e., scale and zero-point) based on the impact of weights on the model output. Experiments on LLaMA and LLaMA-2 show that DAQ consistently outperforms the best baseline method, reducing perplexity loss by an average of 22.8% on LLaMA and 19.6% on LLaMA-2. Our code is available at https://github.com/LuoYingSong/DAQ.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12187
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
Luo, Yingsong
Chen, Ling
Machine Learning
Artificial Intelligence
Large language models (LLMs) excel in various tasks but face deployment challenges due to hardware constraints. We propose density-aware post-training weight-only quantization (DAQ), which has two stages: 1) density-centric alignment, which identifies the center of high-density weights and centers the dynamic range on this point to align high-density weight regions with floating-point high-precision regions; 2) learnable dynamic range adjustment, which adjusts the dynamic range by optimizing quantization parameters (i.e., scale and zero-point) based on the impact of weights on the model output. Experiments on LLaMA and LLaMA-2 show that DAQ consistently outperforms the best baseline method, reducing perplexity loss by an average of 22.8% on LLaMA and 19.6% on LLaMA-2. Our code is available at https://github.com/LuoYingSong/DAQ.
title DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.12187