Optimizing GEMM for Energy and Performance on Versal ACAP Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Papalamprou, Ilias, Masouros, Dimosthenis, Loudaros, Ioannis, Catthoor, Francky, Soudris, Dimitrios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908640572604416
author Papalamprou, Ilias
Masouros, Dimosthenis
Loudaros, Ioannis
Catthoor, Francky
Soudris, Dimitrios
author_facet Papalamprou, Ilias
Masouros, Dimosthenis
Loudaros, Ioannis
Catthoor, Francky
Soudris, Dimitrios
contents General Matrix Multiplication (GEMM) is a fundamental operation in many scientific workloads, signal processing, and particularly deep learning. It is often a bottleneck for performance and energy efficiency, especially in edge environments with tight resource and power constraints. AMD's Versal ACAP offers heterogeneous components (AIEs, PL, PS) that can address these challenges, but mapping GEMM across them is complex, with prior works largely overlooking energy-performance trade-offs. In this paper, we propose an automated framework for Versal ACAP that generates GEMM mappings optimized for either performance or energy efficiency. Unlike prior analytical approaches, our method leverages a Machine Learning (ML) model, trained on approximately 6000 on-board experiments of different GEMM mappings, to guide Design Space Exploration, yielding more efficient designs. Evaluation on the Versal VCK190 shows geomean improvements of 1.23x (up to 2.5x) in throughput and 1.25x (up to 2.7x) in energy efficiency over state-of-the-art frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06907
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
Papalamprou, Ilias
Masouros, Dimosthenis
Loudaros, Ioannis
Catthoor, Francky
Soudris, Dimitrios
Hardware Architecture
General Matrix Multiplication (GEMM) is a fundamental operation in many scientific workloads, signal processing, and particularly deep learning. It is often a bottleneck for performance and energy efficiency, especially in edge environments with tight resource and power constraints. AMD's Versal ACAP offers heterogeneous components (AIEs, PL, PS) that can address these challenges, but mapping GEMM across them is complex, with prior works largely overlooking energy-performance trade-offs. In this paper, we propose an automated framework for Versal ACAP that generates GEMM mappings optimized for either performance or energy efficiency. Unlike prior analytical approaches, our method leverages a Machine Learning (ML) model, trained on approximately 6000 on-board experiments of different GEMM mappings, to guide Design Space Exploration, yielding more efficient designs. Evaluation on the Versal VCK190 shows geomean improvements of 1.23x (up to 2.5x) in throughput and 1.25x (up to 2.7x) in energy efficiency over state-of-the-art frameworks.
title Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
topic Hardware Architecture
url https://arxiv.org/abs/2511.06907