Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alexandridis, Kosmas, Peltekis, Christodoulos, Filippas, Dionysios, Dimitrakopoulos, Giorgos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916365572505600
author Alexandridis, Kosmas
Peltekis, Christodoulos
Filippas, Dionysios
Dimitrakopoulos, Giorgos
author_facet Alexandridis, Kosmas
Peltekis, Christodoulos
Filippas, Dionysios
Dimitrakopoulos, Giorgos
contents The widespread adoption of machine learning algorithms necessitates hardware acceleration to ensure efficient performance. This acceleration relies on custom matrix engines that operate on full or reduced-precision floating-point arithmetic. However, conventional floating-point implementations can be power hungry. This paper proposes a method to improve the energy efficiency of the matrix engines used in machine learning algorithm acceleration. Our approach leverages approximate normalization within the floating-point multiply-add units as a means to reduce their hardware complexity, without sacrificing overall machine-learning model accuracy. Hardware synthesis results show that this technique reduces area and power consumption roughly by 16% and 13% on average for Bfloat16 format. Also, the error introduced in transformer model accuracy is 1% on average, for the most efficient configuration of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11997
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
Alexandridis, Kosmas
Peltekis, Christodoulos
Filippas, Dionysios
Dimitrakopoulos, Giorgos
Hardware Architecture
The widespread adoption of machine learning algorithms necessitates hardware acceleration to ensure efficient performance. This acceleration relies on custom matrix engines that operate on full or reduced-precision floating-point arithmetic. However, conventional floating-point implementations can be power hungry. This paper proposes a method to improve the energy efficiency of the matrix engines used in machine learning algorithm acceleration. Our approach leverages approximate normalization within the floating-point multiply-add units as a means to reduce their hardware complexity, without sacrificing overall machine-learning model accuracy. Hardware synthesis results show that this technique reduces area and power consumption roughly by 16% and 13% on average for Bfloat16 format. Also, the error introduced in transformer model accuracy is 1% on average, for the most efficient configuration of the proposed approach.
title Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
topic Hardware Architecture
url https://arxiv.org/abs/2408.11997