Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929412642963456 |
|---|---|
| author | Zhu, Xuqi Zhang, Huaizhi Lee, JunKyu Zhu, Jiacheng Pal, Chandrajit Saha, Sangeet McDonald-Maier, Klaus D. Zhai, Xiaojun |
| author_facet | Zhu, Xuqi Zhang, Huaizhi Lee, JunKyu Zhu, Jiacheng Pal, Chandrajit Saha, Sangeet McDonald-Maier, Klaus D. Zhai, Xiaojun |
| contents | Modern Neural Network (NN) architectures heavily rely on vast numbers of multiply-accumulate arithmetic operations, constituting the predominant computational cost. Therefore, this paper proposes a high-throughput, scalable and energy efficient non-element-wise matrix multiplication unit on FPGAs as a basic component of the NNs. We firstly streamline inter-layer and intra-layer redundancies of MADDNESS algorithm, a LUT-based approximate matrix multiplication, to design a fast, efficient scalable approximate matrix multiplication module termed "Approximate Multiplication Unit (AMU)". The AMU optimizes LUT-based matrix multiplications further through dedicated memory management and access design, decoupling computational overhead from input resolution and boosting FPGA-based NN accelerator efficiency significantly. The experimental results show that using our AMU achieves up to 9x higher throughput and 112x higher energy efficiency over the state-of-the-art solutions for the FPGA-based Quantised Neural Network (QNN) accelerators. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_02362 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA Zhu, Xuqi Zhang, Huaizhi Lee, JunKyu Zhu, Jiacheng Pal, Chandrajit Saha, Sangeet McDonald-Maier, Klaus D. Zhai, Xiaojun Hardware Architecture Artificial Intelligence Machine Learning Modern Neural Network (NN) architectures heavily rely on vast numbers of multiply-accumulate arithmetic operations, constituting the predominant computational cost. Therefore, this paper proposes a high-throughput, scalable and energy efficient non-element-wise matrix multiplication unit on FPGAs as a basic component of the NNs. We firstly streamline inter-layer and intra-layer redundancies of MADDNESS algorithm, a LUT-based approximate matrix multiplication, to design a fast, efficient scalable approximate matrix multiplication module termed "Approximate Multiplication Unit (AMU)". The AMU optimizes LUT-based matrix multiplications further through dedicated memory management and access design, decoupling computational overhead from input resolution and boosting FPGA-based NN accelerator efficiency significantly. The experimental results show that using our AMU achieves up to 9x higher throughput and 112x higher energy efficiency over the state-of-the-art solutions for the FPGA-based Quantised Neural Network (QNN) accelerators. |
| title | Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA |
| topic | Hardware Architecture Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2407.02362 |