A Power-Efficient Hardware Implementation of L-Mul

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ruiqi, Lyu, Yangxintong, Bao, Han, da Silva, Bruno
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909441179254784
author Chen, Ruiqi
Lyu, Yangxintong
Bao, Han
da Silva, Bruno
author_facet Chen, Ruiqi
Lyu, Yangxintong
Bao, Han
da Silva, Bruno
contents Multiplication is a core operation in modern neural network (NN) computations, contributing significantly to energy consumption. The linear-complexity multiplication (L-Mul) algorithm is specifically proposed as an approximate multiplication method for emerging NN models, such as large language model (LLM), to reduce the energy consumption and computational complexity of multiplications. However, hardware implementation designs for L-Mul have not yet been reported. Additionally, 8-bit floating-point (FP8), as an emerging data format, offers a better dynamic range compared to traditional 8-bit integer (INT8), making it increasingly popular and widely adopted in NN computations. This paper thus presents a power-efficient FPGAbased hardware implementation (approximate FP8 multiplier) for L-Mul. The core computation is implemented using the dynamic reconfigurable lookup tables and carry chains primitives available in AMD Xilinx UltraScale/UltraScale+ technology. The accuracy and resource utilization of the approximate multiplier are evaluated and analyzed. Furthermore, the FP8 approximate multiplier is deployed in the inference phase of representative NN models to validate its effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18948
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Power-Efficient Hardware Implementation of L-Mul
Chen, Ruiqi
Lyu, Yangxintong
Bao, Han
da Silva, Bruno
Hardware Architecture
Multiplication is a core operation in modern neural network (NN) computations, contributing significantly to energy consumption. The linear-complexity multiplication (L-Mul) algorithm is specifically proposed as an approximate multiplication method for emerging NN models, such as large language model (LLM), to reduce the energy consumption and computational complexity of multiplications. However, hardware implementation designs for L-Mul have not yet been reported. Additionally, 8-bit floating-point (FP8), as an emerging data format, offers a better dynamic range compared to traditional 8-bit integer (INT8), making it increasingly popular and widely adopted in NN computations. This paper thus presents a power-efficient FPGAbased hardware implementation (approximate FP8 multiplier) for L-Mul. The core computation is implemented using the dynamic reconfigurable lookup tables and carry chains primitives available in AMD Xilinx UltraScale/UltraScale+ technology. The accuracy and resource utilization of the approximate multiplier are evaluated and analyzed. Furthermore, the FP8 approximate multiplier is deployed in the inference phase of representative NN models to validate its effectiveness.
title A Power-Efficient Hardware Implementation of L-Mul
topic Hardware Architecture
url https://arxiv.org/abs/2412.18948