Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Taka, Endri, Roesti, Andre, Melber, Joseph, Vasireddy, Pranathi, Denolf, Kristof, Marculescu, Diana
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915675884224512
author Taka, Endri
Roesti, Andre
Melber, Joseph
Vasireddy, Pranathi
Denolf, Kristof
Marculescu, Diana
author_facet Taka, Endri
Roesti, Andre
Melber, Joseph
Vasireddy, Pranathi
Denolf, Kristof
Marculescu, Diana
contents The high computational and memory demands of modern deep learning (DL) workloads have led to the development of specialized hardware devices from cloud to edge, such as AMD's Ryzen AI XDNA NPUs. Optimizing general matrix multiplication (GEMM) algorithms for these architectures is critical for improving DL workload performance. To this end, this paper presents a common systematic methodology to optimize GEMM workloads across the two current NPU generations, namely XDNA and XDNA2. Our implementations exploit the unique architectural features of AMD's NPUs and address key performance bottlenecks at the system level. End-to-end performance evaluation across various GEMM sizes demonstrates state-of-the-art throughput of up to 6.76 TOPS (XDNA) and 38.05 TOPS (XDNA2) for 8-bit integer (int8) precision. Similarly, for brain floating-point (bf16) precision, our GEMM implementations attain up to 3.14 TOPS (XDNA) and 14.71 TOPS (XDNA2). This work provides significant insights into key performance aspects of optimizing GEMM workloads on Ryzen AI NPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
Taka, Endri
Roesti, Andre
Melber, Joseph
Vasireddy, Pranathi
Denolf, Kristof
Marculescu, Diana
Hardware Architecture
The high computational and memory demands of modern deep learning (DL) workloads have led to the development of specialized hardware devices from cloud to edge, such as AMD's Ryzen AI XDNA NPUs. Optimizing general matrix multiplication (GEMM) algorithms for these architectures is critical for improving DL workload performance. To this end, this paper presents a common systematic methodology to optimize GEMM workloads across the two current NPU generations, namely XDNA and XDNA2. Our implementations exploit the unique architectural features of AMD's NPUs and address key performance bottlenecks at the system level. End-to-end performance evaluation across various GEMM sizes demonstrates state-of-the-art throughput of up to 6.76 TOPS (XDNA) and 38.05 TOPS (XDNA2) for 8-bit integer (int8) precision. Similarly, for brain floating-point (bf16) precision, our GEMM implementations attain up to 3.14 TOPS (XDNA) and 14.71 TOPS (XDNA2). This work provides significant insights into key performance aspects of optimizing GEMM workloads on Ryzen AI NPUs.
title Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
topic Hardware Architecture
url https://arxiv.org/abs/2512.13282