Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Walter, Dominik, Halm, Marita, Seidel, Daniel, Ghosh, Indrayudh, Heidorn, Christian, Hannig, Frank, Teich, Jürgen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912234177822720
author Walter, Dominik
Halm, Marita
Seidel, Daniel
Ghosh, Indrayudh
Heidorn, Christian
Hannig, Frank
Teich, Jürgen
author_facet Walter, Dominik
Halm, Marita
Seidel, Daniel
Ghosh, Indrayudh
Heidorn, Christian
Hannig, Frank
Teich, Jürgen
contents Increasing demands for computing power also propel the need for energy-efficient SoC accelerator architectures. One class of such accelerators are so-called processor arrays, which typically integrate a two-dimensional mesh of interconnected processing elements~(PEs). Such arrays are specifically designed to accelerate the execution of multidimensional nested loops by exploiting the intrinsic parallelism of loops. Moreover, for mapping a given loop nest application, two opposed mapping methods have emerged: Operation-centric and iteration-centric. Both differ in the granularity of the mapping. The operation-centric approach maps individual operations to the PEs of the array, while the iteration-centric approach maps entire tiles of iterations to each PE. The operation-centric approach is applied predominantly for processor arrays often referred to as Coarse-Grained Reconfigurable Arrays~(CGRAs), while processor arrays supporting an iteration-centric approach are referred to as Tightly-Coupled Processor Arrays~(TCPAs) in the following. This work provides a comprehensive comparison of both approaches and related architectures by evaluating their respective benefits and trade-offs. ...
format Preprint
id arxiv_https___arxiv_org_abs_2502_12062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
Walter, Dominik
Halm, Marita
Seidel, Daniel
Ghosh, Indrayudh
Heidorn, Christian
Hannig, Frank
Teich, Jürgen
Hardware Architecture
Increasing demands for computing power also propel the need for energy-efficient SoC accelerator architectures. One class of such accelerators are so-called processor arrays, which typically integrate a two-dimensional mesh of interconnected processing elements~(PEs). Such arrays are specifically designed to accelerate the execution of multidimensional nested loops by exploiting the intrinsic parallelism of loops. Moreover, for mapping a given loop nest application, two opposed mapping methods have emerged: Operation-centric and iteration-centric. Both differ in the granularity of the mapping. The operation-centric approach maps individual operations to the PEs of the array, while the iteration-centric approach maps entire tiles of iterations to each PE. The operation-centric approach is applied predominantly for processor arrays often referred to as Coarse-Grained Reconfigurable Arrays~(CGRAs), while processor arrays supporting an iteration-centric approach are referred to as Tightly-Coupled Processor Arrays~(TCPAs) in the following. This work provides a comprehensive comparison of both approaches and related architectures by evaluating their respective benefits and trade-offs. ...
title Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
topic Hardware Architecture
url https://arxiv.org/abs/2502.12062