Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuxuan, Belda, María José, Castro, Fernando, Olcoz, Katzalin, Atienza, David, Ansaloni, Giovanni
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910172079718400
author Wang, Yuxuan
Belda, María José
Castro, Fernando
Olcoz, Katzalin
Atienza, David
Ansaloni, Giovanni
author_facet Wang, Yuxuan
Belda, María José
Castro, Fernando
Olcoz, Katzalin
Atienza, David
Ansaloni, Giovanni
contents Modern computing workloads commonly involve matrix-matrix multiplication (mmul) as a core computing pattern. Coarse-Grained Reconfigurable Arrays (CGRAs) can flexibly and efficiently support it, since they combine operation-level reconfigurability and high energy efficiency. However, mapping computational kernels that include mmul with state-of-the-art compilation strategies often leads to suboptimal results, since its multi-dimensional structure hampers the uncovering of its inherent parallelism and, ultimately, runtime performance. Here, we take a different position: we introduce a specialized mmul CGRA kernel schedule, parametrizable across different CGRA sizes. Then, we describe a novel compilation methodology that adapts program representations to effectively leverage it, employing polyhedral transformations to analyze complex computational patterns and expose hidden mmul operations through loop reordering and splitting. The identified patterns are then substituted with optimized assembly, while the remaining program sections are compiled independently. CGRA configurations are then generated, encompassing pre-compiled and compiled parts. Our strategy maximizes resource utilization and ultimately run-time performance, even when mmul is not directly apparent in the source code. The experimental results show speedups up to 9.1x across different benchmarks that contain hidden mmuls and CGRA instances of various sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22297
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
Wang, Yuxuan
Belda, María José
Castro, Fernando
Olcoz, Katzalin
Atienza, David
Ansaloni, Giovanni
Hardware Architecture
Modern computing workloads commonly involve matrix-matrix multiplication (mmul) as a core computing pattern. Coarse-Grained Reconfigurable Arrays (CGRAs) can flexibly and efficiently support it, since they combine operation-level reconfigurability and high energy efficiency. However, mapping computational kernels that include mmul with state-of-the-art compilation strategies often leads to suboptimal results, since its multi-dimensional structure hampers the uncovering of its inherent parallelism and, ultimately, runtime performance. Here, we take a different position: we introduce a specialized mmul CGRA kernel schedule, parametrizable across different CGRA sizes. Then, we describe a novel compilation methodology that adapts program representations to effectively leverage it, employing polyhedral transformations to analyze complex computational patterns and expose hidden mmul operations through loop reordering and splitting. The identified patterns are then substituted with optimized assembly, while the remaining program sections are compiled independently. CGRA configurations are then generated, encompassing pre-compiled and compiled parts. Our strategy maximizes resource utilization and ultimately run-time performance, even when mmul is not directly apparent in the source code. The experimental results show speedups up to 9.1x across different benchmarks that contain hidden mmuls and CGRA instances of various sizes.
title Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
topic Hardware Architecture
url https://arxiv.org/abs/2604.22297