Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nayan, Md Mizanur Rahaman, Raj, Ritik, Shaik, Gouse Basha, Krishna, Tushar, Naeemi, Azad J
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910779680227328
author Nayan, Md Mizanur Rahaman
Raj, Ritik
Shaik, Gouse Basha
Krishna, Tushar
Naeemi, Azad J
author_facet Nayan, Md Mizanur Rahaman
Raj, Ritik
Shaik, Gouse Basha
Krishna, Tushar
Naeemi, Azad J
contents General matrix multiplication (GeMM) is a core operation in virtually all AI applications. Systolic array (SA) based architectures have shown great promise as GeMM hardware accelerators thanks to their speed and energy efficiency. Unfortunately, SAs incur a linear delay in filling the operands, due to unidirectional propagation via pipeline latches. In this work, we propose a novel in-array data orchestration technique in SAs where we enable data feeding on the principal diagonal followed by bi-directional propagation. This improves the runtime by up to 2X at minimal hardware overhead. In addition, the proposed data orchestration enables convolution lowering (known as im2col) using a simple hardware support to fully exploit input feature map reuse opportunity and significantly lower the off-chip memory traffic resulting in 1.2X throughput improvement and 2.17X inference energy reduction during YOLOv3 and RESNET50 workload on average. In contrast, conventional data orchestration would require more elaborate hardware and control signals to implement im2col in hardware because of the data skew. We have synthesized and conducted place and route for 16X16 systolic arrays based on the novel and conventional orchestrations using ASAP 7nm PDK and found that our proposed approach results in 0.211% area and 1.6% power overheads.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06043
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col
Nayan, Md Mizanur Rahaman
Raj, Ritik
Shaik, Gouse Basha
Krishna, Tushar
Naeemi, Azad J
Hardware Architecture
General matrix multiplication (GeMM) is a core operation in virtually all AI applications. Systolic array (SA) based architectures have shown great promise as GeMM hardware accelerators thanks to their speed and energy efficiency. Unfortunately, SAs incur a linear delay in filling the operands, due to unidirectional propagation via pipeline latches. In this work, we propose a novel in-array data orchestration technique in SAs where we enable data feeding on the principal diagonal followed by bi-directional propagation. This improves the runtime by up to 2X at minimal hardware overhead. In addition, the proposed data orchestration enables convolution lowering (known as im2col) using a simple hardware support to fully exploit input feature map reuse opportunity and significantly lower the off-chip memory traffic resulting in 1.2X throughput improvement and 2.17X inference energy reduction during YOLOv3 and RESNET50 workload on average. In contrast, conventional data orchestration would require more elaborate hardware and control signals to implement im2col in hardware because of the data skew. We have synthesized and conducted place and route for 16X16 systolic arrays based on the novel and conventional orchestrations using ASAP 7nm PDK and found that our proposed approach results in 0.211% area and 1.6% power overheads.
title Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col
topic Hardware Architecture
url https://arxiv.org/abs/2501.06043