Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Aggarwal, Shivam, Damsgaard, Hans Jakob, Pappalardo, Alessandro, Franco, Giuseppe, Preußer, Thomas B., Blott, Michaela, Mitra, Tulika
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913416728281088
author Aggarwal, Shivam
Damsgaard, Hans Jakob
Pappalardo, Alessandro
Franco, Giuseppe
Preußer, Thomas B.
Blott, Michaela
Mitra, Tulika
author_facet Aggarwal, Shivam
Damsgaard, Hans Jakob
Pappalardo, Alessandro
Franco, Giuseppe
Preußer, Thomas B.
Blott, Michaela
Mitra, Tulika
contents Post-training quantization (PTQ) is a powerful technique for model compression, reducing the numerical precision in neural networks without additional training overhead. Recent works have investigated adopting 8-bit floating-point formats(FP8) in the context of PTQ for model inference. However, floating-point formats smaller than 8 bits and their relative comparison in terms of accuracy-hardware cost with integers remains unexplored on FPGAs. In this work, we present minifloats, which are reduced-precision floating-point formats capable of further reducing the memory footprint, latency, and energy cost of a model while approaching full-precision model accuracy. We implement a custom FPGA-based multiply-accumulate operator library and explore the vast design space, comparing minifloat and integer representations across 3 to 8 bits for both weights and activations. We also examine the applicability of various integerbased quantization techniques to minifloats. Our experiments show that minifloats offer a promising alternative for emerging workloads such as vision transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12359
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
Aggarwal, Shivam
Damsgaard, Hans Jakob
Pappalardo, Alessandro
Franco, Giuseppe
Preußer, Thomas B.
Blott, Michaela
Mitra, Tulika
Computer Vision and Pattern Recognition
Artificial Intelligence
Hardware Architecture
Machine Learning
Performance
Post-training quantization (PTQ) is a powerful technique for model compression, reducing the numerical precision in neural networks without additional training overhead. Recent works have investigated adopting 8-bit floating-point formats(FP8) in the context of PTQ for model inference. However, floating-point formats smaller than 8 bits and their relative comparison in terms of accuracy-hardware cost with integers remains unexplored on FPGAs. In this work, we present minifloats, which are reduced-precision floating-point formats capable of further reducing the memory footprint, latency, and energy cost of a model while approaching full-precision model accuracy. We implement a custom FPGA-based multiply-accumulate operator library and explore the vast design space, comparing minifloat and integer representations across 3 to 8 bits for both weights and activations. We also examine the applicability of various integerbased quantization techniques to minifloats. Our experiments show that minifloats offer a promising alternative for emerging workloads such as vision transformers.
title Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Hardware Architecture
Machine Learning
Performance
url https://arxiv.org/abs/2311.12359