Magic for the Age of Quantized DNNs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sawada, Yoshihide, Saiin, Ryuji, Suetake, Kazuma
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909146561904640
author Sawada, Yoshihide
Saiin, Ryuji
Suetake, Kazuma
author_facet Sawada, Yoshihide
Saiin, Ryuji
Suetake, Kazuma
contents Recently, the number of parameters in DNNs has explosively increased, as exemplified by LLMs (Large Language Models), making inference on small-scale computers more difficult. Model compression technology is, therefore, essential for integration into products. In this paper, we propose a method of quantization-aware training. We introduce a novel normalization (Layer-Batch Normalization) that is independent of the mini-batch size and does not require any additional computation cost during inference. Then, we quantize the weights by the scaled round-clip function with the weight standardization. We also quantize activation functions using the same function and apply surrogate gradients to train the model with both quantized weights and the quantized activation functions. We call this method Magic for the age of Quantised DNNs (MaQD). Experimental results show that our quantization method can be achieved with minimal accuracy degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14999
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Magic for the Age of Quantized DNNs
Sawada, Yoshihide
Saiin, Ryuji
Suetake, Kazuma
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Neural and Evolutionary Computing
Recently, the number of parameters in DNNs has explosively increased, as exemplified by LLMs (Large Language Models), making inference on small-scale computers more difficult. Model compression technology is, therefore, essential for integration into products. In this paper, we propose a method of quantization-aware training. We introduce a novel normalization (Layer-Batch Normalization) that is independent of the mini-batch size and does not require any additional computation cost during inference. Then, we quantize the weights by the scaled round-clip function with the weight standardization. We also quantize activation functions using the same function and apply surrogate gradients to train the model with both quantized weights and the quantized activation functions. We call this method Magic for the age of Quantised DNNs (MaQD). Experimental results show that our quantization method can be achieved with minimal accuracy degradation.
title Magic for the Age of Quantized DNNs
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Neural and Evolutionary Computing
url https://arxiv.org/abs/2403.14999