DR-CGRA: Supporting Loop-Carried Dependencies in CGRAs Without Spilling Intermediate Values
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hadar, Elad, Etsion, Yoav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
von: Walter, Dominik, et al.
Veröffentlicht: (2025)
von: Walter, Dominik, et al.
Veröffentlicht: (2025)
Evaluation of CGRA Toolchains
von: Walter, Dominik, et al.
Veröffentlicht: (2025)
von: Walter, Dominik, et al.
Veröffentlicht: (2025)
Building an Open CGRA Ecosystem for Agile Innovation
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
SAT-based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
von: Tirelli, Cristian, et al.
Veröffentlicht: (2024)
von: Tirelli, Cristian, et al.
Veröffentlicht: (2024)
A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
von: Alexandris, Georgios, et al.
Veröffentlicht: (2025)
von: Alexandris, Georgios, et al.
Veröffentlicht: (2025)
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
von: Vazquez, Daniel, et al.
Veröffentlicht: (2024)
von: Vazquez, Daniel, et al.
Veröffentlicht: (2024)
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
von: Li, Zhaoying, et al.
Veröffentlicht: (2024)
von: Li, Zhaoying, et al.
Veröffentlicht: (2024)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
von: Carpentieri, Nicolò, et al.
Veröffentlicht: (2024)
von: Carpentieri, Nicolò, et al.
Veröffentlicht: (2024)
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
Monomorphism-based CGRA Mapping via Space and Time Decoupling
von: Tirelli, Cristian, et al.
Veröffentlicht: (2025)
von: Tirelli, Cristian, et al.
Veröffentlicht: (2025)
NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
von: Prasad, Rohit
Veröffentlicht: (2025)
von: Prasad, Rohit
Veröffentlicht: (2025)
An ultra-low-power CGRA for accelerating Transformers at the edge
von: Prasad, Rohit
Veröffentlicht: (2025)
von: Prasad, Rohit
Veröffentlicht: (2025)
Fast Bipartitioned Hybrid Adder Utilizing Carry Select and Carry Lookahead Logic
von: Balasubramanian, Padmanabhan, et al.
Veröffentlicht: (2024)
von: Balasubramanian, Padmanabhan, et al.
Veröffentlicht: (2024)
Re-thinking Memory-Bound Limitations in CGRAs
von: Liu, Xiangfeng, et al.
Veröffentlicht: (2025)
von: Liu, Xiangfeng, et al.
Veröffentlicht: (2025)
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
von: Deshmukh, Sameer, et al.
Veröffentlicht: (2024)
von: Deshmukh, Sameer, et al.
Veröffentlicht: (2024)
Dynamic Loop Fusion in High-Level Synthesis
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Closed-Loop Environmental Control System on Embedded Systems
von: Goswami, Irisha M., et al.
Veröffentlicht: (2026)
von: Goswami, Irisha M., et al.
Veröffentlicht: (2026)
Mestra: Exploring Migration on Virtualized CGRAs
von: Kyriazis, Agamemnon, et al.
Veröffentlicht: (2026)
von: Kyriazis, Agamemnon, et al.
Veröffentlicht: (2026)
Symbolic Polyhedral-Based Energy Analysis for Nested Loop Programs
von: Nirmala, Avinash Mahesh, et al.
Veröffentlicht: (2026)
von: Nirmala, Avinash Mahesh, et al.
Veröffentlicht: (2026)
Loop Control Management in Tightly Coupled Processor Arrays (TCPAs)
von: Walter, Dominik, et al.
Veröffentlicht: (2026)
von: Walter, Dominik, et al.
Veröffentlicht: (2026)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025)
von: Ren, Yi, et al.
Veröffentlicht: (2025)
Verification and Validation (V&V)-in-the-Loop for RISC-V Design: The Holistic Vision of BZL
von: Ahmed, Sajjad, et al.
Veröffentlicht: (2026)
von: Ahmed, Sajjad, et al.
Veröffentlicht: (2026)
A Logarithmic Depth Quantum Carry-Lookahead Modulo $(2^n-1)$ Adder
von: Gaur, Bhaskar, et al.
Veröffentlicht: (2024)
von: Gaur, Bhaskar, et al.
Veröffentlicht: (2024)
Adding MFMA Support to gem5
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Demystifying the 7-D Convolution Loop Nest for Data and Instruction Streaming in Reconfigurable AI Accelerators
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2025)
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2025)
Extend IVerilog to Support Batch RTL Fault Simulation
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
Support Vector Machines Classification on Bendable RISC-V
von: Vergos, Polykarpos, et al.
Veröffentlicht: (2025)
von: Vergos, Polykarpos, et al.
Veröffentlicht: (2025)
Multiport Support for Vortex OpenGPU Memory Hierarchy
von: Shin, Injae, et al.
Veröffentlicht: (2025)
von: Shin, Injae, et al.
Veröffentlicht: (2025)
ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
von: Durvasula, Sankeerth, et al.
Veröffentlicht: (2024)
von: Durvasula, Sankeerth, et al.
Veröffentlicht: (2024)
An Optimal Alignment-Driven Iterative Closed-Loop Convergence Framework for High-Performance Ultra-Large Scale Layout Pattern Clustering
von: Liu, Shuo
Veröffentlicht: (2025)
von: Liu, Shuo
Veröffentlicht: (2025)
THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX
von: Dizani, Farshad, et al.
Veröffentlicht: (2025)
von: Dizani, Farshad, et al.
Veröffentlicht: (2025)
3D MPSoC with On-Chip Cache Support -- Design and Exploitation
von: Cataldo, Rodrigo, et al.
Veröffentlicht: (2025)
von: Cataldo, Rodrigo, et al.
Veröffentlicht: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
von: Ding, Hong, et al.
Veröffentlicht: (2025)
von: Ding, Hong, et al.
Veröffentlicht: (2025)
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025)
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025)
VIKIN: A Reconfigurable Accelerator for KANs and MLPs with Two-Stage Sparsity Support
von: Ou, Wenhui, et al.
Veröffentlicht: (2026)
von: Ou, Wenhui, et al.
Veröffentlicht: (2026)
VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling
von: Lin, Zi-Wei, et al.
Veröffentlicht: (2026)
von: Lin, Zi-Wei, et al.
Veröffentlicht: (2026)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
von: Langarita, Rubén, et al.
Veröffentlicht: (2025)
von: Langarita, Rubén, et al.
Veröffentlicht: (2025)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
von: Walter, Dominik, et al.
Veröffentlicht: (2025) -
Evaluation of CGRA Toolchains
von: Walter, Dominik, et al.
Veröffentlicht: (2025) -
Building an Open CGRA Ecosystem for Agile Innovation
von: Juneja, Rohan, et al.
Veröffentlicht: (2025) -
SAT-based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
von: Tirelli, Cristian, et al.
Veröffentlicht: (2024) -
A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
von: Alexandris, Georgios, et al.
Veröffentlicht: (2025)