MSQ: Memory-Efficient Bit Sparsification Quantization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Han, Seokho, Yoon, Seoyeon, Kim, Jinhee, Wang, Dongwei, Jeon, Kang Eun, Yang, Huanrui, Ko, Jong Hwan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909712125001728
author Han, Seokho
Yoon, Seoyeon
Kim, Jinhee
Wang, Dongwei
Jeon, Kang Eun
Yang, Huanrui
Ko, Jong Hwan
author_facet Han, Seokho
Yoon, Seoyeon
Kim, Jinhee
Wang, Dongwei
Jeon, Kang Eun
Yang, Huanrui
Ko, Jong Hwan
contents As deep neural networks (DNNs) see increased deployment on mobile and edge devices, optimizing model efficiency has become crucial. Mixed-precision quantization is widely favored, as it offers a superior balance between efficiency and accuracy compared to uniform quantization. However, finding the optimal precision for each layer is challenging. Recent studies utilizing bit-level sparsity have shown promise, yet they often introduce substantial training complexity and high GPU memory requirements. In this paper, we propose Memory-Efficient Bit Sparsification Quantization (MSQ), a novel approach that addresses these limitations. MSQ applies a round-clamp quantizer to enable differentiable computation of the least significant bits (LSBs) from model weights. It further employs regularization to induce sparsity in these LSBs, enabling effective precision reduction without explicit bit-level parameter splitting. Additionally, MSQ incorporates Hessian information, allowing the simultaneous pruning of multiple LSBs to further enhance training efficiency. Experimental results show that MSQ achieves up to 8.00x reduction in trainable parameters and up to 86% reduction in training time compared to previous bit-level quantization, while maintaining competitive accuracy and compression rates. This makes it a practical solution for training efficient DNNs on resource-constrained devices.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MSQ: Memory-Efficient Bit Sparsification Quantization
Han, Seokho
Yoon, Seoyeon
Kim, Jinhee
Wang, Dongwei
Jeon, Kang Eun
Yang, Huanrui
Ko, Jong Hwan
Machine Learning
As deep neural networks (DNNs) see increased deployment on mobile and edge devices, optimizing model efficiency has become crucial. Mixed-precision quantization is widely favored, as it offers a superior balance between efficiency and accuracy compared to uniform quantization. However, finding the optimal precision for each layer is challenging. Recent studies utilizing bit-level sparsity have shown promise, yet they often introduce substantial training complexity and high GPU memory requirements. In this paper, we propose Memory-Efficient Bit Sparsification Quantization (MSQ), a novel approach that addresses these limitations. MSQ applies a round-clamp quantizer to enable differentiable computation of the least significant bits (LSBs) from model weights. It further employs regularization to induce sparsity in these LSBs, enabling effective precision reduction without explicit bit-level parameter splitting. Additionally, MSQ incorporates Hessian information, allowing the simultaneous pruning of multiple LSBs to further enhance training efficiency. Experimental results show that MSQ achieves up to 8.00x reduction in trainable parameters and up to 86% reduction in training time compared to previous bit-level quantization, while maintaining competitive accuracy and compression rates. This makes it a practical solution for training efficient DNNs on resource-constrained devices.
title MSQ: Memory-Efficient Bit Sparsification Quantization
topic Machine Learning
url https://arxiv.org/abs/2507.22349