Can Asymmetric Tile Buffering Be Beneficial?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909915384119296 |
|---|---|
| author | Wang, Chengyue Pang, Wesley Wu, Xinrui Jun, Gregory Romero, Luis Taka, Endri Marculescu, Diana Nowatzki, Tony Vasireddy, Pranathi Melber, Joseph Chen, Deming Cong, Jason |
| author_facet | Wang, Chengyue Pang, Wesley Wu, Xinrui Jun, Gregory Romero, Luis Taka, Endri Marculescu, Diana Nowatzki, Tony Vasireddy, Pranathi Melber, Joseph Chen, Deming Cong, Jason |
| contents | General matrix multiplication (GEMM) is the computational backbone of modern AI workloads, and its efficiency is critically dependent on effective tiling strategies. Conventional approaches employ symmetric tile buffering, where the buffered tile size of the input $A$ along the dimension $M$ matches the output tile size of $C$.
In this paper, we introduce asymmetric tile buffering (ATB), a simple but powerful technique that decouples the buffered tile dimensions of the input and output operands. We show, for the first time, that ATB is both practical and highly beneficial. To explain this effect, we develop a performance model that incorporates both the benefits of ATB (higher arithmetic intensity) and its overheads (higher kernel switching costs), providing insight into how to select effective ATB tiling factors. As a case study, we apply ATB to AMD's latest XDNA2 AI Engine (AIE), achieving up to a 4.54x speedup, from 4.8 to 24.6 TFLOPS on mixed-precision BFP16--BF16 GEMM, establishing a new performance record for XDNA2 AIE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_16041 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Can Asymmetric Tile Buffering Be Beneficial? Wang, Chengyue Pang, Wesley Wu, Xinrui Jun, Gregory Romero, Luis Taka, Endri Marculescu, Diana Nowatzki, Tony Vasireddy, Pranathi Melber, Joseph Chen, Deming Cong, Jason Distributed, Parallel, and Cluster Computing Hardware Architecture Performance General matrix multiplication (GEMM) is the computational backbone of modern AI workloads, and its efficiency is critically dependent on effective tiling strategies. Conventional approaches employ symmetric tile buffering, where the buffered tile size of the input $A$ along the dimension $M$ matches the output tile size of $C$. In this paper, we introduce asymmetric tile buffering (ATB), a simple but powerful technique that decouples the buffered tile dimensions of the input and output operands. We show, for the first time, that ATB is both practical and highly beneficial. To explain this effect, we develop a performance model that incorporates both the benefits of ATB (higher arithmetic intensity) and its overheads (higher kernel switching costs), providing insight into how to select effective ATB tiling factors. As a case study, we apply ATB to AMD's latest XDNA2 AI Engine (AIE), achieving up to a 4.54x speedup, from 4.8 to 24.6 TFLOPS on mixed-precision BFP16--BF16 GEMM, establishing a new performance record for XDNA2 AIE. |
| title | Can Asymmetric Tile Buffering Be Beneficial? |
| topic | Distributed, Parallel, and Cluster Computing Hardware Architecture Performance |
| url | https://arxiv.org/abs/2511.16041 |