Quantization-Aware Regularizers for Deep Neural Networks Compression

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Malchiodi, Dario, Ferraretto, Mattia, Frasca, Marco
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914304400293888
author Malchiodi, Dario
Ferraretto, Mattia
Frasca, Marco
author_facet Malchiodi, Dario
Ferraretto, Mattia
Frasca, Marco
contents Deep Neural Networks reached state-of-the-art performance across numerous domains, but this progress has come at the cost of increasingly large and over-parameterized models, posing serious challenges for deployment on resource-constrained devices. As a result, model compression has become essential, and -- among compression techniques -- weight quantization is largely used and particularly effective, yet it typically introduces a non-negligible accuracy drop. However, it is usually applied to already trained models, without influencing how the parameter space is explored during the learning phase. In contrast, we introduce per-layer regularization terms that drive weights to naturally form clusters during training, integrating quantization awareness directly into the optimization process. This reduces the accuracy loss typically associated with quantization methods while preserving their compression potential. Furthermore, in our framework quantization representatives become network parameters, marking, to the best of our knowledge, the first approach to embed quantization parameters directly into the backpropagation procedure. Experiments on CIFAR-10 with AlexNet and VGG16 models confirm the effectiveness of the proposed strategy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03614
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantization-Aware Regularizers for Deep Neural Networks Compression
Malchiodi, Dario
Ferraretto, Mattia
Frasca, Marco
Machine Learning
Deep Neural Networks reached state-of-the-art performance across numerous domains, but this progress has come at the cost of increasingly large and over-parameterized models, posing serious challenges for deployment on resource-constrained devices. As a result, model compression has become essential, and -- among compression techniques -- weight quantization is largely used and particularly effective, yet it typically introduces a non-negligible accuracy drop. However, it is usually applied to already trained models, without influencing how the parameter space is explored during the learning phase. In contrast, we introduce per-layer regularization terms that drive weights to naturally form clusters during training, integrating quantization awareness directly into the optimization process. This reduces the accuracy loss typically associated with quantization methods while preserving their compression potential. Furthermore, in our framework quantization representatives become network parameters, marking, to the best of our knowledge, the first approach to embed quantization parameters directly into the backpropagation procedure. Experiments on CIFAR-10 with AlexNet and VGG16 models confirm the effectiveness of the proposed strategy.
title Quantization-Aware Regularizers for Deep Neural Networks Compression
topic Machine Learning
url https://arxiv.org/abs/2602.03614