Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhaoying, Dangi, Pranav, Yin, Chenyang, Bandara, Thilini Kaushalya, Juneja, Rohan, Tan, Cheng, Bai, Zhenyu, Mitra, Tulika |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Building an Open CGRA Ecosystem for Agile Innovation
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
Sustainable Hardware Specialization
by: Dangi, Pranav, et al.
Published: (2024)
by: Dangi, Pranav, et al.
Published: (2024)
A Data-Driven Dynamic Execution Orchestration Architecture
by: Bai, Zhenyu, et al.
Published: (2026)
by: Bai, Zhenyu, et al.
Published: (2026)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025)
by: Yin, Chenyang, et al.
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
by: Bai, Zhenyu, et al.
Published: (2025)
by: Bai, Zhenyu, et al.
Published: (2025)
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
Evaluation of CGRA Toolchains
by: Walter, Dominik, et al.
Published: (2025)
by: Walter, Dominik, et al.
Published: (2025)
Hybrid Photonic-digital Accelerator for Attention Mechanism
by: Li, Huize, et al.
Published: (2025)
by: Li, Huize, et al.
Published: (2025)
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
by: Vazquez, Daniel, et al.
Published: (2024)
by: Vazquez, Daniel, et al.
Published: (2024)
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
by: Wang, Yuxuan, et al.
Published: (2026)
by: Wang, Yuxuan, et al.
Published: (2026)
Monomorphism-based CGRA Mapping via Space and Time Decoupling
by: Tirelli, Cristian, et al.
Published: (2025)
by: Tirelli, Cristian, et al.
Published: (2025)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
by: Carpentieri, Nicolò, et al.
Published: (2024)
by: Carpentieri, Nicolò, et al.
Published: (2024)
NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
DR-CGRA: Supporting Loop-Carried Dependencies in CGRAs Without Spilling Intermediate Values
by: Hadar, Elad, et al.
Published: (2024)
by: Hadar, Elad, et al.
Published: (2024)
SparrowSNN: A Hardware/software Co-design for Energy Efficient ECG Classification
by: Yan, Zhanglu, et al.
Published: (2024)
by: Yan, Zhanglu, et al.
Published: (2024)
An ultra-low-power CGRA for accelerating Transformers at the edge
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026)
by: Zhao, Liang, et al.
Published: (2026)
Time Reversal for Near-Field Communications on Multi-chip Wireless Networks
by: Rodríguez-Galán, Fátima, et al.
Published: (2024)
by: Rodríguez-Galán, Fátima, et al.
Published: (2024)
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
by: Ai, Chenyang, et al.
Published: (2024)
by: Ai, Chenyang, et al.
Published: (2024)
Enhanced Hybrid Temporal Computing Using Deterministic Summations for Ultra-Low-Power Accelerators
by: Sachdeva, Sachin, et al.
Published: (2025)
by: Sachdeva, Sachin, et al.
Published: (2025)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
by: Ai, Chenyang, et al.
Published: (2026)
by: Ai, Chenyang, et al.
Published: (2026)
SRAM Based Digital Custom Compute Engine for Improved Area Efficiency of AI Hardware
by: Dhakad, Narendra Singh, et al.
Published: (2026)
by: Dhakad, Narendra Singh, et al.
Published: (2026)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
by: Perotti, Matteo, et al.
Published: (2023)
by: Perotti, Matteo, et al.
Published: (2023)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
In-Storage Domain-Specific Acceleration for Serverless Computing
by: Mahapatra, Rohan, et al.
Published: (2023)
by: Mahapatra, Rohan, et al.
Published: (2023)
Enhancing Computational Efficiency in Intensive Domains via Redundant Residue Number Systems
by: Mousavi, Soudabeh, et al.
Published: (2024)
by: Mousavi, Soudabeh, et al.
Published: (2024)
Modeling PFAS in Semiconductor Manufacturing to Quantify Trade-offs in Energy Efficiency and Environmental Impact of Computing Systems
by: Elgamal, Mariam, et al.
Published: (2025)
by: Elgamal, Mariam, et al.
Published: (2025)
UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM Architecture
by: Chen, Sitian, et al.
Published: (2024)
by: Chen, Sitian, et al.
Published: (2024)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
by: Tan, Zhanhong, et al.
Published: (2024)
by: Tan, Zhanhong, et al.
Published: (2024)
Accelerating Electrostatics-based Global Placement with Enhanced FFT Computation
by: Zhang, Hangyu, et al.
Published: (2025)
by: Zhang, Hangyu, et al.
Published: (2025)
A Fully Pipelined FIFO Based Polynomial Multiplication Hardware Architecture Based On Number Theoretic Transform
by: Heidarpur, Moslem, et al.
Published: (2025)
by: Heidarpur, Moslem, et al.
Published: (2025)
HaVen: Hallucination-Mitigated LLM for Verilog Code Generation Aligned with HDL Engineers
by: Yang, Yiyao, et al.
Published: (2025)
by: Yang, Yiyao, et al.
Published: (2025)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
by: Ma, Siyuan, et al.
Published: (2025)
by: Ma, Siyuan, et al.
Published: (2025)
ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic Interpolation
by: Deng, Chenhui, et al.
Published: (2024)
by: Deng, Chenhui, et al.
Published: (2024)
AutoVeriFix: Automatically Correcting Errors and Enhancing Functional Correctness in LLM-Generated Verilog Code
by: Tan, Yan, et al.
Published: (2025)
by: Tan, Yan, et al.
Published: (2025)
A Novel Cost-Effective MIMO Architecture with Ray Antenna Array for Enhanced Wireless Communication Performance
by: Dong, Zhenjun, et al.
Published: (2025)
by: Dong, Zhenjun, et al.
Published: (2025)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
by: Lin, Ye, et al.
Published: (2026)
by: Lin, Ye, et al.
Published: (2026)
Similar Items
-
Building an Open CGRA Ecosystem for Agile Innovation
by: Juneja, Rohan, et al.
Published: (2025) -
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
by: Juneja, Rohan, et al.
Published: (2025) -
Sustainable Hardware Specialization
by: Dangi, Pranav, et al.
Published: (2024) -
A Data-Driven Dynamic Execution Orchestration Architecture
by: Bai, Zhenyu, et al.
Published: (2026) -
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025)