FPGA Resource-aware Structured Pruning for Real-Time Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramhorst, Benjamin, Loncar, Vladimir, Constantinides, George A.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917988950605824
author Ramhorst, Benjamin
Loncar, Vladimir
Constantinides, George A.
author_facet Ramhorst, Benjamin
Loncar, Vladimir
Constantinides, George A.
contents Neural networks achieve state-of-the-art performance in image classification, speech recognition, scientific analysis and many more application areas. Due to the high computational complexity and memory footprint of neural networks, various compression techniques, such as pruning and quantization, have been proposed in literature. Pruning sparsifies a neural network, reducing the number of multiplications and memory. However, pruning often fails to capture properties of the underlying hardware, causing unstructured sparsity and load-balance inefficiency, thus bottlenecking resource improvements. We propose a hardware-centric formulation of pruning, by formulating it as a knapsack problem with resource-aware tensor structures. Evaluated on a range of tasks, including sub-microsecond particle classification at CERN's Large Hadron Collider and fast image classification, the proposed method achieves reductions ranging between 55% and 92% in the DSP utilization and up to 81% in BRAM utilization.
format Preprint
id arxiv_https___arxiv_org_abs_2308_05170
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
Ramhorst, Benjamin
Loncar, Vladimir
Constantinides, George A.
Hardware Architecture
Artificial Intelligence
Neural networks achieve state-of-the-art performance in image classification, speech recognition, scientific analysis and many more application areas. Due to the high computational complexity and memory footprint of neural networks, various compression techniques, such as pruning and quantization, have been proposed in literature. Pruning sparsifies a neural network, reducing the number of multiplications and memory. However, pruning often fails to capture properties of the underlying hardware, causing unstructured sparsity and load-balance inefficiency, thus bottlenecking resource improvements. We propose a hardware-centric formulation of pruning, by formulating it as a knapsack problem with resource-aware tensor structures. Evaluated on a range of tasks, including sub-microsecond particle classification at CERN's Large Hadron Collider and fast image classification, the proposed method achieves reductions ranging between 55% and 92% in the DSP utilization and up to 81% in BRAM utilization.
title FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
topic Hardware Architecture
Artificial Intelligence
url https://arxiv.org/abs/2308.05170