LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yanyue, Li, Zhengang, Diaconu, Dana, Handagala, Suranga, Leeser, Miriam, Lin, Xue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910703992963072
author Xie, Yanyue
Li, Zhengang
Diaconu, Dana
Handagala, Suranga
Leeser, Miriam
Lin, Xue
author_facet Xie, Yanyue
Li, Zhengang
Diaconu, Dana
Handagala, Suranga
Leeser, Miriam
Lin, Xue
contents For FPGA-based neural network accelerators, digital signal processing (DSP) blocks have traditionally been the cornerstone for handling multiplications. This paper introduces LUTMUL, which harnesses the potential of look-up tables (LUTs) for performing multiplications. The availability of LUTs typically outnumbers that of DSPs by a factor of 100, offering a significant computational advantage. By exploiting this advantage of LUTs, our method demonstrates a potential boost in the performance of FPGA-based neural network accelerators with a reconfigurable dataflow architecture. Our approach challenges the conventional peak performance on DSP-based accelerators and sets a new benchmark for efficient neural network inference on FPGAs. Experimental results demonstrate that our design achieves the best inference speed among all FPGA-based accelerators, achieving a throughput of 1627 images per second and maintaining a top-1 accuracy of 70.95% on the ImageNet dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11852
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
Xie, Yanyue
Li, Zhengang
Diaconu, Dana
Handagala, Suranga
Leeser, Miriam
Lin, Xue
Hardware Architecture
Artificial Intelligence
Machine Learning
For FPGA-based neural network accelerators, digital signal processing (DSP) blocks have traditionally been the cornerstone for handling multiplications. This paper introduces LUTMUL, which harnesses the potential of look-up tables (LUTs) for performing multiplications. The availability of LUTs typically outnumbers that of DSPs by a factor of 100, offering a significant computational advantage. By exploiting this advantage of LUTs, our method demonstrates a potential boost in the performance of FPGA-based neural network accelerators with a reconfigurable dataflow architecture. Our approach challenges the conventional peak performance on DSP-based accelerators and sets a new benchmark for efficient neural network inference on FPGAs. Experimental results demonstrate that our design achieves the best inference speed among all FPGA-based accelerators, achieving a throughput of 1627 images per second and maintaining a top-1 accuracy of 70.95% on the ImageNet dataset.
title LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
topic Hardware Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.11852