FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Müller, Mika Markus, Lübeck, Konstantin, Jung, Alexander Louis-Ferdinand, Steinmetz, Jannik, Bringmann, Oliver
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915318232776704
author Müller, Mika Markus
Lübeck, Konstantin
Jung, Alexander Louis-Ferdinand
Steinmetz, Jannik
Bringmann, Oliver
author_facet Müller, Mika Markus
Lübeck, Konstantin
Jung, Alexander Louis-Ferdinand
Steinmetz, Jannik
Bringmann, Oliver
contents Artificial Intelligence (AI) algorithms, such as Deep Neural Networks (DNNs), have become an important tool for a wide range of applications, from computer vision to natural language processing. However, the computational complexity of DNN inference poses a significant challenge, particularly for processing on resource-constrained edge devices. One promising approach to address this challenge is the exploitation of sparsity in DNN operator weights. In this work, we present FlexiSAGA, an architecturally configurable and dataflow-flexible AI hardware accelerator for the sparse and dense processing of general matrix multiplications (GEMMs). FlexiSAGA supports seven different sparse and dense dataflows, enabling efficient processing of resource intensive DNN operators. Additionally, we propose a DNN pruning method specifically tailored towards the FlexiSAGA architecture, allowing for near-optimal processing of dense and sparse convolution and fully-connected operators, facilitating a DNN/HW co-design flow. Our results show a whole DNN sparse-over-dense inference speedup ranging from 1.41 up to 4.28, outperforming commercial and literature-reported accelerator platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01566
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
Müller, Mika Markus
Lübeck, Konstantin
Jung, Alexander Louis-Ferdinand
Steinmetz, Jannik
Bringmann, Oliver
Performance
Artificial Intelligence
Hardware Architecture
Machine Learning
Artificial Intelligence (AI) algorithms, such as Deep Neural Networks (DNNs), have become an important tool for a wide range of applications, from computer vision to natural language processing. However, the computational complexity of DNN inference poses a significant challenge, particularly for processing on resource-constrained edge devices. One promising approach to address this challenge is the exploitation of sparsity in DNN operator weights. In this work, we present FlexiSAGA, an architecturally configurable and dataflow-flexible AI hardware accelerator for the sparse and dense processing of general matrix multiplications (GEMMs). FlexiSAGA supports seven different sparse and dense dataflows, enabling efficient processing of resource intensive DNN operators. Additionally, we propose a DNN pruning method specifically tailored towards the FlexiSAGA architecture, allowing for near-optimal processing of dense and sparse convolution and fully-connected operators, facilitating a DNN/HW co-design flow. Our results show a whole DNN sparse-over-dense inference speedup ranging from 1.41 up to 4.28, outperforming commercial and literature-reported accelerator platforms.
title FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
topic Performance
Artificial Intelligence
Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2506.01566