Soft Quantization: Model Compression Via Weight Coupling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bernstein, Daniel T., Di Carlo, Luca, Schwab, David
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917230924529664
author Bernstein, Daniel T.
Di Carlo, Luca
Schwab, David
author_facet Bernstein, Daniel T.
Di Carlo, Luca
Schwab, David
contents We show that introducing short-range attractive couplings between the weights of a neural network during training provides a novel avenue for model quantization. These couplings rapidly induce the discretization of a model's weight distribution, and they do so in a mixed-precision manner despite only relying on two additional hyperparameters. We demonstrate that, within an appropriate range of hyperparameters, our "soft quantization'' scheme outperforms histogram-equalized post-training quantization on ResNet-20/CIFAR-10. Soft quantization provides both a new pipeline for the flexible compression of machine learning models and a new tool for investigating the trade-off between compression and generalization in high-dimensional loss landscapes.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21219
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Soft Quantization: Model Compression Via Weight Coupling
Bernstein, Daniel T.
Di Carlo, Luca
Schwab, David
Machine Learning
Disordered Systems and Neural Networks
We show that introducing short-range attractive couplings between the weights of a neural network during training provides a novel avenue for model quantization. These couplings rapidly induce the discretization of a model's weight distribution, and they do so in a mixed-precision manner despite only relying on two additional hyperparameters. We demonstrate that, within an appropriate range of hyperparameters, our "soft quantization'' scheme outperforms histogram-equalized post-training quantization on ResNet-20/CIFAR-10. Soft quantization provides both a new pipeline for the flexible compression of machine learning models and a new tool for investigating the trade-off between compression and generalization in high-dimensional loss landscapes.
title Soft Quantization: Model Compression Via Weight Coupling
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2601.21219