Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Xuqi, Zhang, Huaizhi, Lee, JunKyu, Zhu, Jiacheng, Pal, Chandrajit, Saha, Sangeet, McDonald-Maier, Klaus D., Zhai, Xiaojun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929412642963456
author Zhu, Xuqi
Zhang, Huaizhi
Lee, JunKyu
Zhu, Jiacheng
Pal, Chandrajit
Saha, Sangeet
McDonald-Maier, Klaus D.
Zhai, Xiaojun
author_facet Zhu, Xuqi
Zhang, Huaizhi
Lee, JunKyu
Zhu, Jiacheng
Pal, Chandrajit
Saha, Sangeet
McDonald-Maier, Klaus D.
Zhai, Xiaojun
contents Modern Neural Network (NN) architectures heavily rely on vast numbers of multiply-accumulate arithmetic operations, constituting the predominant computational cost. Therefore, this paper proposes a high-throughput, scalable and energy efficient non-element-wise matrix multiplication unit on FPGAs as a basic component of the NNs. We firstly streamline inter-layer and intra-layer redundancies of MADDNESS algorithm, a LUT-based approximate matrix multiplication, to design a fast, efficient scalable approximate matrix multiplication module termed "Approximate Multiplication Unit (AMU)". The AMU optimizes LUT-based matrix multiplications further through dedicated memory management and access design, decoupling computational overhead from input resolution and boosting FPGA-based NN accelerator efficiency significantly. The experimental results show that using our AMU achieves up to 9x higher throughput and 112x higher energy efficiency over the state-of-the-art solutions for the FPGA-based Quantised Neural Network (QNN) accelerators.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02362
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
Zhu, Xuqi
Zhang, Huaizhi
Lee, JunKyu
Zhu, Jiacheng
Pal, Chandrajit
Saha, Sangeet
McDonald-Maier, Klaus D.
Zhai, Xiaojun
Hardware Architecture
Artificial Intelligence
Machine Learning
Modern Neural Network (NN) architectures heavily rely on vast numbers of multiply-accumulate arithmetic operations, constituting the predominant computational cost. Therefore, this paper proposes a high-throughput, scalable and energy efficient non-element-wise matrix multiplication unit on FPGAs as a basic component of the NNs. We firstly streamline inter-layer and intra-layer redundancies of MADDNESS algorithm, a LUT-based approximate matrix multiplication, to design a fast, efficient scalable approximate matrix multiplication module termed "Approximate Multiplication Unit (AMU)". The AMU optimizes LUT-based matrix multiplications further through dedicated memory management and access design, decoupling computational overhead from input resolution and boosting FPGA-based NN accelerator efficiency significantly. The experimental results show that using our AMU achieves up to 9x higher throughput and 112x higher energy efficiency over the state-of-the-art solutions for the FPGA-based Quantised Neural Network (QNN) accelerators.
title Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
topic Hardware Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.02362