Evaluation of CGRA Toolchains
Fuente:
arXiv
Guardado en:
| Autores principales: | Walter, Dominik, Halm, Marita, Seidel, Daniel, Ghosh, Indrayudh, Heidorn, Christian, Hannig, Frank, Teich, Jürgen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
por: Walter, Dominik, et al.
Publicado: (2025)
por: Walter, Dominik, et al.
Publicado: (2025)
Loop Control Management in Tightly Coupled Processor Arrays (TCPAs)
por: Walter, Dominik, et al.
Publicado: (2026)
por: Walter, Dominik, et al.
Publicado: (2026)
Symbolic Polyhedral-Based Energy Analysis for Nested Loop Programs
por: Nirmala, Avinash Mahesh, et al.
Publicado: (2026)
por: Nirmala, Avinash Mahesh, et al.
Publicado: (2026)
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
por: Morales, José Juan Hernández, et al.
Publicado: (2026)
por: Morales, José Juan Hernández, et al.
Publicado: (2026)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
por: Sabih, Muhammad, et al.
Publicado: (2025)
por: Sabih, Muhammad, et al.
Publicado: (2025)
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
por: Vazquez, Daniel, et al.
Publicado: (2024)
por: Vazquez, Daniel, et al.
Publicado: (2024)
Building an Open CGRA Ecosystem for Agile Innovation
por: Juneja, Rohan, et al.
Publicado: (2025)
por: Juneja, Rohan, et al.
Publicado: (2025)
Monomorphism-based CGRA Mapping via Space and Time Decoupling
por: Tirelli, Cristian, et al.
Publicado: (2025)
por: Tirelli, Cristian, et al.
Publicado: (2025)
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
por: Wang, Yuxuan, et al.
Publicado: (2026)
por: Wang, Yuxuan, et al.
Publicado: (2026)
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
por: Li, Zhaoying, et al.
Publicado: (2024)
por: Li, Zhaoying, et al.
Publicado: (2024)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
por: Carpentieri, Nicolò, et al.
Publicado: (2024)
por: Carpentieri, Nicolò, et al.
Publicado: (2024)
NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
por: Prasad, Rohit
Publicado: (2025)
por: Prasad, Rohit
Publicado: (2025)
DR-CGRA: Supporting Loop-Carried Dependencies in CGRAs Without Spilling Intermediate Values
por: Hadar, Elad, et al.
Publicado: (2024)
por: Hadar, Elad, et al.
Publicado: (2024)
An ultra-low-power CGRA for accelerating Transformers at the edge
por: Prasad, Rohit
Publicado: (2025)
por: Prasad, Rohit
Publicado: (2025)
From PyTorch to Calyx: An Open-Source Compiler Toolchain for ML Accelerators
por: Xie, Jiahan, et al.
Publicado: (2025)
por: Xie, Jiahan, et al.
Publicado: (2025)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
por: Iff, Patrick, et al.
Publicado: (2023)
por: Iff, Patrick, et al.
Publicado: (2023)
OpTC -- A Toolchain for Deployment of Neural Networks on AURIX TC3xx Microcontrollers
por: Heidorn, Christian, et al.
Publicado: (2024)
por: Heidorn, Christian, et al.
Publicado: (2024)
A High-level Synthesis Toolchain for the Julia Language
por: Short, Benedict, et al.
Publicado: (2025)
por: Short, Benedict, et al.
Publicado: (2025)
Modeling and Simulating Emerging Memory Technologies: A Tutorial
por: Chen, Yun-Chih, et al.
Publicado: (2025)
por: Chen, Yun-Chih, et al.
Publicado: (2025)
The Impact of Logic Locking on Confidentiality: An Automated Evaluation
por: Reimann, Lennart M., et al.
Publicado: (2025)
por: Reimann, Lennart M., et al.
Publicado: (2025)
Realizing Hardware-Optimized General Tree-Based Data Structures for Heterogeneous System Classes
por: Biebert, Daniel, et al.
Publicado: (2025)
por: Biebert, Daniel, et al.
Publicado: (2025)
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
por: Deshmukh, Sameer, et al.
Publicado: (2024)
por: Deshmukh, Sameer, et al.
Publicado: (2024)
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
por: Ghosh, Arkapravo, et al.
Publicado: (2025)
por: Ghosh, Arkapravo, et al.
Publicado: (2025)
A RISC-V MCU with adaptive reverse body bias and ultra-low-power retention mode in 22 nm FD-SOI
por: Bauer, Heiner, et al.
Publicado: (2023)
por: Bauer, Heiner, et al.
Publicado: (2023)
Work-in-Progress: Real-Time Neural Network Inference on a Custom RISC-V Multicore Vector Processor
por: Kirschner, Maximilian, et al.
Publicado: (2024)
por: Kirschner, Maximilian, et al.
Publicado: (2024)
Hardware and software build flow with SoCMake
por: Pejašinović, Risto, et al.
Publicado: (2025)
por: Pejašinović, Risto, et al.
Publicado: (2025)
SA-DS: A Dataset for Large Language Model-Driven AI Accelerator Design Generation
por: Vungarala, Deepak, et al.
Publicado: (2024)
por: Vungarala, Deepak, et al.
Publicado: (2024)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
por: Su, Yuchao, et al.
Publicado: (2025)
por: Su, Yuchao, et al.
Publicado: (2025)
COmPOSER: Circuit Optimization of mm-wave/RF circuits with Performance-Oriented Synthesis for Efficient Realizations
por: Ghosh, Subhadip, et al.
Publicado: (2026)
por: Ghosh, Subhadip, et al.
Publicado: (2026)
MultiVic: A Time-Predictable RISC-V Multi-Core Processor Optimized for Neural Network Inference
por: Kirschner, Maximilian, et al.
Publicado: (2025)
por: Kirschner, Maximilian, et al.
Publicado: (2025)
Toward designing workload-aware Surface Code Architectures
por: Ghosh, Archisman, et al.
Publicado: (2026)
por: Ghosh, Archisman, et al.
Publicado: (2026)
No Tile Left Behind: Multiprogramming for Surface-Code Architectures
por: Ghosh, Archisman, et al.
Publicado: (2026)
por: Ghosh, Archisman, et al.
Publicado: (2026)
Belenos: Bottleneck Evaluation to Link Biomechanics to Novel Computing Optimizations
por: Chitsaz, Hana, et al.
Publicado: (2025)
por: Chitsaz, Hana, et al.
Publicado: (2025)
Further Evaluations of a Didactic CPU Visual Simulator (CPUVSIM)
por: Cortinovis, Renato, et al.
Publicado: (2024)
por: Cortinovis, Renato, et al.
Publicado: (2024)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2023)
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2023)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
por: Koenig, Cyril, et al.
Publicado: (2025)
por: Koenig, Cyril, et al.
Publicado: (2025)
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
por: Parvathy, Aradhana Mohan, et al.
Publicado: (2026)
por: Parvathy, Aradhana Mohan, et al.
Publicado: (2026)
Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator
por: Ando, Takuto, et al.
Publicado: (2025)
por: Ando, Takuto, et al.
Publicado: (2025)
Evaluation of NVENC Split-Frame Encoding (SFE) for UHD Video Transcoding
por: Arunruangsirilert, Kasidis, et al.
Publicado: (2025)
por: Arunruangsirilert, Kasidis, et al.
Publicado: (2025)
Efficient Trace for RISC-V: Design, Evaluation, and Integration in CVA6
por: Laghi, Umberto, et al.
Publicado: (2025)
por: Laghi, Umberto, et al.
Publicado: (2025)
Ejemplares similares
-
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
por: Walter, Dominik, et al.
Publicado: (2025) -
Loop Control Management in Tightly Coupled Processor Arrays (TCPAs)
por: Walter, Dominik, et al.
Publicado: (2026) -
Symbolic Polyhedral-Based Energy Analysis for Nested Loop Programs
por: Nirmala, Avinash Mahesh, et al.
Publicado: (2026) -
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
por: Morales, José Juan Hernández, et al.
Publicado: (2026) -
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
por: Sabih, Muhammad, et al.
Publicado: (2025)