PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lou, Binglei, Rademacher, Richard, Boland, David, Leong, Philip H. W.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917775406006272
author Lou, Binglei
Rademacher, Richard
Boland, David
Leong, Philip H. W.
author_facet Lou, Binglei
Rademacher, Richard
Boland, David
Leong, Philip H. W.
contents FPGAs have distinct advantages as a technology for deploying deep neural networks (DNNs) at the edge. Lookup Table (LUT) based networks, where neurons are directly modeled using LUTs, help maximize this promise of offering ultra-low latency and high area efficiency on FPGAs. Unfortunately, LUT resource usage scales exponentially with the number of inputs to the LUT, restricting PolyLUT to small LUT sizes. This work introduces PolyLUT-Add, a technique that enhances neuron connectivity by combining $A$ PolyLUT sub-neurons via addition to improve accuracy. Moreover, we describe a novel architecture to improve its scalability. We evaluated our implementation over the MNIST, Jet Substructure classification, and Network Intrusion Detection benchmark and found that for similar accuracy, PolyLUT-Add achieves a LUT reduction of $2.0-13.9\times$ with a $1.2-1.6\times$ decrease in latency.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04910
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
Lou, Binglei
Rademacher, Richard
Boland, David
Leong, Philip H. W.
Machine Learning
Artificial Intelligence
Hardware Architecture
FPGAs have distinct advantages as a technology for deploying deep neural networks (DNNs) at the edge. Lookup Table (LUT) based networks, where neurons are directly modeled using LUTs, help maximize this promise of offering ultra-low latency and high area efficiency on FPGAs. Unfortunately, LUT resource usage scales exponentially with the number of inputs to the LUT, restricting PolyLUT to small LUT sizes. This work introduces PolyLUT-Add, a technique that enhances neuron connectivity by combining $A$ PolyLUT sub-neurons via addition to improve accuracy. Moreover, we describe a novel architecture to improve its scalability. We evaluated our implementation over the MNIST, Jet Substructure classification, and Network Intrusion Detection benchmark and found that for similar accuracy, PolyLUT-Add achieves a LUT reduction of $2.0-13.9\times$ with a $1.2-1.6\times$ decrease in latency.
title PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
topic Machine Learning
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2406.04910