Scaling Law for Quantization-Aware Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Mengzhao, Zhang, Chaoyi, Liu, Jing, Zeng, Yutao, Xue, Zeyue, Liu, Zhiheng, Li, Yunshui, Ma, Jin, Huang, Jie, Zhou, Xun, Luo, Ping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916746220273664
author Chen, Mengzhao
Zhang, Chaoyi
Liu, Jing
Zeng, Yutao
Xue, Zeyue
Liu, Zhiheng
Li, Yunshui
Ma, Jin
Huang, Jie
Zhou, Xun
Luo, Ping
author_facet Chen, Mengzhao
Zhang, Chaoyi
Liu, Jing
Zeng, Yutao
Xue, Zeyue
Liu, Zhiheng
Li, Yunshui
Ma, Jin
Huang, Jie
Zhou, Xun
Luo, Ping
contents Large language models (LLMs) demand substantial computational and memory resources, creating deployment challenges. Quantization-aware training (QAT) addresses these challenges by reducing model precision while maintaining performance. However, the scaling behavior of QAT, especially at 4-bit precision (W4A4), is not well understood. Existing QAT scaling laws often ignore key factors such as the number of training tokens and quantization granularity, which limits their applicability. This paper proposes a unified scaling law for QAT that models quantization error as a function of model size, training data volume, and quantization group size. Through 268 QAT experiments, we show that quantization error decreases as model size increases, but rises with more training tokens and coarser quantization granularity. To identify the sources of W4A4 quantization error, we decompose it into weight and activation components. Both components follow the overall trend of W4A4 quantization error, but with different sensitivities. Specifically, weight quantization error increases more rapidly with more training tokens. Further analysis shows that the activation quantization error in the FC2 layer, caused by outliers, is the primary bottleneck of W4A4 QAT quantization error. By applying mixed-precision quantization to address this bottleneck, we demonstrate that weight and activation quantization errors can converge to similar levels. Additionally, with more training data, weight quantization error eventually exceeds activation quantization error, suggesting that reducing weight quantization error is also important in such scenarios. These findings offer key insights for improving QAT research and development.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Law for Quantization-Aware Training
Chen, Mengzhao
Zhang, Chaoyi
Liu, Jing
Zeng, Yutao
Xue, Zeyue
Liu, Zhiheng
Li, Yunshui
Ma, Jin
Huang, Jie
Zhou, Xun
Luo, Ping
Machine Learning
Computation and Language
Large language models (LLMs) demand substantial computational and memory resources, creating deployment challenges. Quantization-aware training (QAT) addresses these challenges by reducing model precision while maintaining performance. However, the scaling behavior of QAT, especially at 4-bit precision (W4A4), is not well understood. Existing QAT scaling laws often ignore key factors such as the number of training tokens and quantization granularity, which limits their applicability. This paper proposes a unified scaling law for QAT that models quantization error as a function of model size, training data volume, and quantization group size. Through 268 QAT experiments, we show that quantization error decreases as model size increases, but rises with more training tokens and coarser quantization granularity. To identify the sources of W4A4 quantization error, we decompose it into weight and activation components. Both components follow the overall trend of W4A4 quantization error, but with different sensitivities. Specifically, weight quantization error increases more rapidly with more training tokens. Further analysis shows that the activation quantization error in the FC2 layer, caused by outliers, is the primary bottleneck of W4A4 QAT quantization error. By applying mixed-precision quantization to address this bottleneck, we demonstrate that weight and activation quantization errors can converge to similar levels. Additionally, with more training data, weight quantization error eventually exceeds activation quantization error, suggesting that reducing weight quantization error is also important in such scenarios. These findings offer key insights for improving QAT research and development.
title Scaling Law for Quantization-Aware Training
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.14302