Can Asymmetric Tile Buffering Be Beneficial?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chengyue, Pang, Wesley, Wu, Xinrui, Jun, Gregory, Romero, Luis, Taka, Endri, Marculescu, Diana, Nowatzki, Tony, Vasireddy, Pranathi, Melber, Joseph, Chen, Deming, Cong, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909915384119296
author Wang, Chengyue
Pang, Wesley
Wu, Xinrui
Jun, Gregory
Romero, Luis
Taka, Endri
Marculescu, Diana
Nowatzki, Tony
Vasireddy, Pranathi
Melber, Joseph
Chen, Deming
Cong, Jason
author_facet Wang, Chengyue
Pang, Wesley
Wu, Xinrui
Jun, Gregory
Romero, Luis
Taka, Endri
Marculescu, Diana
Nowatzki, Tony
Vasireddy, Pranathi
Melber, Joseph
Chen, Deming
Cong, Jason
contents General matrix multiplication (GEMM) is the computational backbone of modern AI workloads, and its efficiency is critically dependent on effective tiling strategies. Conventional approaches employ symmetric tile buffering, where the buffered tile size of the input $A$ along the dimension $M$ matches the output tile size of $C$. In this paper, we introduce asymmetric tile buffering (ATB), a simple but powerful technique that decouples the buffered tile dimensions of the input and output operands. We show, for the first time, that ATB is both practical and highly beneficial. To explain this effect, we develop a performance model that incorporates both the benefits of ATB (higher arithmetic intensity) and its overheads (higher kernel switching costs), providing insight into how to select effective ATB tiling factors. As a case study, we apply ATB to AMD's latest XDNA2 AI Engine (AIE), achieving up to a 4.54x speedup, from 4.8 to 24.6 TFLOPS on mixed-precision BFP16--BF16 GEMM, establishing a new performance record for XDNA2 AIE.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16041
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Asymmetric Tile Buffering Be Beneficial?
Wang, Chengyue
Pang, Wesley
Wu, Xinrui
Jun, Gregory
Romero, Luis
Taka, Endri
Marculescu, Diana
Nowatzki, Tony
Vasireddy, Pranathi
Melber, Joseph
Chen, Deming
Cong, Jason
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Performance
General matrix multiplication (GEMM) is the computational backbone of modern AI workloads, and its efficiency is critically dependent on effective tiling strategies. Conventional approaches employ symmetric tile buffering, where the buffered tile size of the input $A$ along the dimension $M$ matches the output tile size of $C$. In this paper, we introduce asymmetric tile buffering (ATB), a simple but powerful technique that decouples the buffered tile dimensions of the input and output operands. We show, for the first time, that ATB is both practical and highly beneficial. To explain this effect, we develop a performance model that incorporates both the benefits of ATB (higher arithmetic intensity) and its overheads (higher kernel switching costs), providing insight into how to select effective ATB tiling factors. As a case study, we apply ATB to AMD's latest XDNA2 AI Engine (AIE), achieving up to a 4.54x speedup, from 4.8 to 24.6 TFLOPS on mixed-precision BFP16--BF16 GEMM, establishing a new performance record for XDNA2 AIE.
title Can Asymmetric Tile Buffering Be Beneficial?
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
Performance
url https://arxiv.org/abs/2511.16041