A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Alexandris, Georgios, Chaidos, Panagiotis, Maras, Alexis, de Bruin, Barry, Gomony, Manil Dev, Corporaal, Henk, Soudris, Dimitrios, Xydis, Sotirios
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916766532239360
author Alexandris, Georgios
Chaidos, Panagiotis
Maras, Alexis
de Bruin, Barry
Gomony, Manil Dev
Corporaal, Henk
Soudris, Dimitrios
Xydis, Sotirios
author_facet Alexandris, Georgios
Chaidos, Panagiotis
Maras, Alexis
de Bruin, Barry
Gomony, Manil Dev
Corporaal, Henk
Soudris, Dimitrios
Xydis, Sotirios
contents The ever-increasing complexity and operational diversity of modern Neural Networks (NNs) have caused the need for low-power and, at the same time, high-performance edge devices for AI applications. Coarse Grained Reconfigurable Architectures (CGRAs) form a promising design paradigm to address these challenges, delivering a close-to-ASIC performance while allowing for hardware programmability. In this paper, we introduce a novel end-to-end exploration and synthesis framework for approximate CGRA processors that enables transparent and optimized integration and mapping of state-of-the-art approximate multiplication components into CGRAs. Our methodology introduces a per-channel exploration strategy that maps specific output features onto approximate components based on accuracy degradation constraints. This enables the optimization of the system's energy consumption while retaining the accuracy above a certain threshold. At the circuit level, the integration of approximate components enables the creation of voltage islands that operate at reduced voltage levels, which is attributed to their inherently shorter critical paths. This key enabler allows us to effectively reduce the overall power consumption by an average of 30% across our analyzed architectures, compared to their baseline counterparts, while incurring only a minimal 2% area overhead. The proposed methodology was evaluated on a widely used NN model, MobileNetV2, on the ImageNet dataset, demonstrating that the generated architectures can deliver up to 440 GOPS/W with relatively small output error during inference, outperforming several State-of-the-Art CGRA architectures in terms of throughput and energy efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
Alexandris, Georgios
Chaidos, Panagiotis
Maras, Alexis
de Bruin, Barry
Gomony, Manil Dev
Corporaal, Henk
Soudris, Dimitrios
Xydis, Sotirios
Hardware Architecture
The ever-increasing complexity and operational diversity of modern Neural Networks (NNs) have caused the need for low-power and, at the same time, high-performance edge devices for AI applications. Coarse Grained Reconfigurable Architectures (CGRAs) form a promising design paradigm to address these challenges, delivering a close-to-ASIC performance while allowing for hardware programmability. In this paper, we introduce a novel end-to-end exploration and synthesis framework for approximate CGRA processors that enables transparent and optimized integration and mapping of state-of-the-art approximate multiplication components into CGRAs. Our methodology introduces a per-channel exploration strategy that maps specific output features onto approximate components based on accuracy degradation constraints. This enables the optimization of the system's energy consumption while retaining the accuracy above a certain threshold. At the circuit level, the integration of approximate components enables the creation of voltage islands that operate at reduced voltage levels, which is attributed to their inherently shorter critical paths. This key enabler allows us to effectively reduce the overall power consumption by an average of 30% across our analyzed architectures, compared to their baseline counterparts, while incurring only a minimal 2% area overhead. The proposed methodology was evaluated on a widely used NN model, MobileNetV2, on the ImageNet dataset, demonstrating that the generated architectures can deliver up to 440 GOPS/W with relatively small output error during inference, outperforming several State-of-the-Art CGRA architectures in terms of throughput and energy efficiency.
title A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
topic Hardware Architecture
url https://arxiv.org/abs/2505.23553