Frame Quantization of Neural Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Czaja, Wojciech, Na, Sanghoon
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909167338389504
author Czaja, Wojciech
Na, Sanghoon
author_facet Czaja, Wojciech
Na, Sanghoon
contents We present a post-training quantization algorithm with error estimates relying on ideas originating from frame theory. Specifically, we use first-order Sigma-Delta ($ΣΔ$) quantization for finite unit-norm tight frames to quantize weight matrices and biases in a neural network. In our scenario, we derive an error bound between the original neural network and the quantized neural network in terms of step size and the number of frame elements. We also demonstrate how to leverage the redundancy of frames to achieve a quantized neural network with higher accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08131
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Frame Quantization of Neural Networks
Czaja, Wojciech
Na, Sanghoon
Machine Learning
Information Theory
We present a post-training quantization algorithm with error estimates relying on ideas originating from frame theory. Specifically, we use first-order Sigma-Delta ($ΣΔ$) quantization for finite unit-norm tight frames to quantize weight matrices and biases in a neural network. In our scenario, we derive an error bound between the original neural network and the quantized neural network in terms of step size and the number of frame elements. We also demonstrate how to leverage the redundancy of frames to achieve a quantized neural network with higher accuracy.
title Frame Quantization of Neural Networks
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2404.08131