Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuxuan, Belda, María José, Castro, Fernando, Olcoz, Katzalin, Atienza, David, Ansaloni, Giovanni |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimization of a Line Detection Algorithm for Autonomous Vehicles on a RISC-V with Accelerator
by: Belda, María José, et al.
Published: (2024)
by: Belda, María José, et al.
Published: (2024)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025)
by: Aspros, Maxime Henri, et al.
Published: (2025)
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
by: Vazquez, Daniel, et al.
Published: (2024)
by: Vazquez, Daniel, et al.
Published: (2024)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
by: Carpentieri, Nicolò, et al.
Published: (2024)
by: Carpentieri, Nicolò, et al.
Published: (2024)
X-HEEP: An Open-Source, Configurable and Extendible RISC-V Platform for TinyAI Applications
by: Machetti, Simone, et al.
Published: (2025)
by: Machetti, Simone, et al.
Published: (2025)
Evaluation of CGRA Toolchains
by: Walter, Dominik, et al.
Published: (2025)
by: Walter, Dominik, et al.
Published: (2025)
e-GPU: An Open-Source and Configurable RISC-V Graphic Processing Unit for TinyAI Applications
by: Machetti, Simone, et al.
Published: (2025)
by: Machetti, Simone, et al.
Published: (2025)
SAT-based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
by: Tirelli, Cristian, et al.
Published: (2024)
by: Tirelli, Cristian, et al.
Published: (2024)
Building an Open CGRA Ecosystem for Agile Innovation
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
by: Li, Zhaoying, et al.
Published: (2024)
by: Li, Zhaoying, et al.
Published: (2024)
Monomorphism-based CGRA Mapping via Space and Time Decoupling
by: Tirelli, Cristian, et al.
Published: (2025)
by: Tirelli, Cristian, et al.
Published: (2025)
Physical Design Exploration of a Wire-Friendly Domain-Specific Processor for Angstrom-Era Nodes
by: Ruotolo, Lorenzo, et al.
Published: (2025)
by: Ruotolo, Lorenzo, et al.
Published: (2025)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026)
by: Villarrubia, Jorge, et al.
Published: (2026)
NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
DR-CGRA: Supporting Loop-Carried Dependencies in CGRAs Without Spilling Intermediate Values
by: Hadar, Elad, et al.
Published: (2024)
by: Hadar, Elad, et al.
Published: (2024)
An ultra-low-power CGRA for accelerating Transformers at the edge
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
MetaWearS: A Shortcut in Wearable Systems Lifecycle with Only a Few Shots
by: Amirshahi, Alireza, et al.
Published: (2024)
by: Amirshahi, Alireza, et al.
Published: (2024)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Increasing the Energy-Efficiency of Wearables Using Low-Precision Posit Arithmetic with PHEE
by: Mallasén, David, et al.
Published: (2025)
by: Mallasén, David, et al.
Published: (2025)
VCO-CARE: VCO-based Calibration-free Analog Readout for Electrodermal activity sensing
by: Alvero-Gonzalez, Leidy Mabel, et al.
Published: (2025)
by: Alvero-Gonzalez, Leidy Mabel, et al.
Published: (2025)
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
by: Palacios, Pedro, et al.
Published: (2024)
by: Palacios, Pedro, et al.
Published: (2024)
CXLRAMSim v1.0: System-Level Exploration of CXL Memory Expander Cards
by: Pathak, Karan, et al.
Published: (2026)
by: Pathak, Karan, et al.
Published: (2026)
X-HEEP: An Open-Source, Configurable and Extendible RISC-V Microcontroller for the Exploration of Ultra-Low-Power Edge Accelerators
by: Machetti, Simone, et al.
Published: (2024)
by: Machetti, Simone, et al.
Published: (2024)
Invited Paper: FEMU: An Open-Source and Configurable Emulation Framework for Prototyping TinyAI Heterogeneous Systems
by: Machetti, Simone, et al.
Published: (2025)
by: Machetti, Simone, et al.
Published: (2025)
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024)
by: Wang, Luming, et al.
Published: (2024)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
by: Caon, Michele, et al.
Published: (2024)
by: Caon, Michele, et al.
Published: (2024)
Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications
by: Cammarata, Danilo, et al.
Published: (2025)
by: Cammarata, Danilo, et al.
Published: (2025)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
by: Duan, Cenlin, et al.
Published: (2024)
by: Duan, Cenlin, et al.
Published: (2024)
Teaching Experiences using the RVfpga Package
by: Chaver, D., et al.
Published: (2024)
by: Chaver, D., et al.
Published: (2024)
Just TestIt! An SBST Approach To Automate System-Integration Testing
by: Terzano, Tommaso, et al.
Published: (2025)
by: Terzano, Tommaso, et al.
Published: (2025)
ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses
by: Li, Mengming, et al.
Published: (2026)
by: Li, Mengming, et al.
Published: (2026)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
QUADOL: A Quality-Driven Approximate Logic Synthesis Method Exploiting Dual-Output LUTs for Modern FPGAs
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
by: Espinosa, Elena, et al.
Published: (2025)
by: Espinosa, Elena, et al.
Published: (2025)
L-PCN: A Point Cloud Accelerator Exploiting Spatial Locality through Octree-based Islandization
by: Gao, Yiming, et al.
Published: (2026)
by: Gao, Yiming, et al.
Published: (2026)
Exploiting Control-flow Enforcement Technology for Sound and Precise Static Binary Disassembly
by: Zhao, Brian, et al.
Published: (2025)
by: Zhao, Brian, et al.
Published: (2025)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
by: Ko, Sho, et al.
Published: (2024)
by: Ko, Sho, et al.
Published: (2024)
Similar Items
-
Optimization of a Line Detection Algorithm for Autonomous Vehicles on a RISC-V with Accelerator
by: Belda, María José, et al.
Published: (2024) -
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025) -
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
by: Vazquez, Daniel, et al.
Published: (2024) -
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
by: Carpentieri, Nicolò, et al.
Published: (2024) -
X-HEEP: An Open-Source, Configurable and Extendible RISC-V Platform for TinyAI Applications
by: Machetti, Simone, et al.
Published: (2025)