How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: You, Jaeseong, Park, Minseop, Lee, Kyunggeun, An, Seokjun, Patel, Chirag, Nage, Markus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917651230490624
author You, Jaeseong
Park, Minseop
Lee, Kyunggeun
An, Seokjun
Patel, Chirag
Nage, Markus
author_facet You, Jaeseong
Park, Minseop
Lee, Kyunggeun
An, Seokjun
Patel, Chirag
Nage, Markus
contents This paper investigates three different parameterizations of asymmetric uniform quantization for quantization-aware training: (1) scale and offset, (2) minimum and maximum, and (3) beta and gamma. We perform a comprehensive comparative analysis of these parameterizations' influence on quantization-aware training, using both controlled experiments and real-world large language models. Our particular focus is on their changing behavior in response to critical training hyperparameters, bit width and learning rate. Based on our investigation, we propose best practices to stabilize and accelerate quantization-aware training with learnable asymmetric quantization ranges.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16898
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
You, Jaeseong
Park, Minseop
Lee, Kyunggeun
An, Seokjun
Patel, Chirag
Nage, Markus
Machine Learning
Artificial Intelligence
This paper investigates three different parameterizations of asymmetric uniform quantization for quantization-aware training: (1) scale and offset, (2) minimum and maximum, and (3) beta and gamma. We perform a comprehensive comparative analysis of these parameterizations' influence on quantization-aware training, using both controlled experiments and real-world large language models. Our particular focus is on their changing behavior in response to critical training hyperparameters, bit width and learning rate. Based on our investigation, we propose best practices to stabilize and accelerate quantization-aware training with learnable asymmetric quantization ranges.
title How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2404.16898