Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhaoying, Dangi, Pranav, Yin, Chenyang, Bandara, Thilini Kaushalya, Juneja, Rohan, Tan, Cheng, Bai, Zhenyu, Mitra, Tulika
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915060438269952
author Li, Zhaoying
Dangi, Pranav
Yin, Chenyang
Bandara, Thilini Kaushalya
Juneja, Rohan
Tan, Cheng
Bai, Zhenyu
Mitra, Tulika
author_facet Li, Zhaoying
Dangi, Pranav
Yin, Chenyang
Bandara, Thilini Kaushalya
Juneja, Rohan
Tan, Cheng
Bai, Zhenyu
Mitra, Tulika
contents Coarse-grained Reconfigurable Arrays (CGRAs) are domain-agnostic accelerators that enhance the energy efficiency of resource-constrained edge devices. The CGRA landscape is diverse, exhibiting trade-offs between performance, efficiency, and architectural specialization. However, CGRAs often overprovision communication resources relative to their modest computing capabilities. This occurs because the theoretically provisioned programmability for CGRAs often proves superfluous in practical implementations. In this paper, we propose Plaid, a novel CGRA architecture and compiler that aligns compute and communication capabilities, thereby significantly improving energy and area efficiency while preserving its generality and performance. We demonstrate that the dataflow graph, representing the target application, can be decomposed into smaller, recurring communication patterns called motifs. The primary contribution is the identification of these structural motifs within the dataflow graphs and the development of an efficient collective execution and routing strategy tailored to these motifs. The Plaid architecture employs a novel collective processing unit that can execute multiple operations of a motif and route related data dependencies together. The Plaid compiler can hierarchically map the dataflow graph and judiciously schedule the motifs. Our design achieves a 43% reduction in power consumption and 46% area savings compared to the baseline high-performance spatio-temporal CGRA, all while preserving its generality and performance levels. In comparison to the baseline energy-efficient spatial CGRA, Plaid offers a 1.4x performance improvement and a 48% area savings, with almost the same power.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08137
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
Li, Zhaoying
Dangi, Pranav
Yin, Chenyang
Bandara, Thilini Kaushalya
Juneja, Rohan
Tan, Cheng
Bai, Zhenyu
Mitra, Tulika
Hardware Architecture
Coarse-grained Reconfigurable Arrays (CGRAs) are domain-agnostic accelerators that enhance the energy efficiency of resource-constrained edge devices. The CGRA landscape is diverse, exhibiting trade-offs between performance, efficiency, and architectural specialization. However, CGRAs often overprovision communication resources relative to their modest computing capabilities. This occurs because the theoretically provisioned programmability for CGRAs often proves superfluous in practical implementations. In this paper, we propose Plaid, a novel CGRA architecture and compiler that aligns compute and communication capabilities, thereby significantly improving energy and area efficiency while preserving its generality and performance. We demonstrate that the dataflow graph, representing the target application, can be decomposed into smaller, recurring communication patterns called motifs. The primary contribution is the identification of these structural motifs within the dataflow graphs and the development of an efficient collective execution and routing strategy tailored to these motifs. The Plaid architecture employs a novel collective processing unit that can execute multiple operations of a motif and route related data dependencies together. The Plaid compiler can hierarchically map the dataflow graph and judiciously schedule the motifs. Our design achieves a 43% reduction in power consumption and 46% area savings compared to the baseline high-performance spatio-temporal CGRA, all while preserving its generality and performance levels. In comparison to the baseline energy-efficient spatial CGRA, Plaid offers a 1.4x performance improvement and a 48% area savings, with almost the same power.
title Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
topic Hardware Architecture
url https://arxiv.org/abs/2412.08137