Re-thinking Memory-Bound Limitations in CGRAs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiangfeng, Jiang, Zhe, Zhu, Anzhen, Han, Xiaomeng, Lyu, Mingsong, Deng, Qingxu, Guan, Nan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Weight/Output Stationarity
by: Chauvaux, Nicolas, et al.
Published: (2024)
by: Chauvaux, Nicolas, et al.
Published: (2024)
Resource Optimized Quantum Squaring Circuit
by: Sultana, Afrin, et al.
Published: (2024)
by: Sultana, Afrin, et al.
Published: (2024)
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
by: Goswami, Kaustav, et al.
Published: (2026)
by: Goswami, Kaustav, et al.
Published: (2026)
Chameleon: A MatMul-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
by: Blanken, Douwe den, et al.
Published: (2025)
by: Blanken, Douwe den, et al.
Published: (2025)
parti-gem5: gem5's Timing Mode Parallelised
by: Cubero-Cascante, José, et al.
Published: (2023)
by: Cubero-Cascante, José, et al.
Published: (2023)
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026)
by: Kyriazis, Agamemnon, et al.
Published: (2026)
Hourglass Sorting: A novel parallel sorting algorithm and its implementation
by: Bascones, Daniel, et al.
Published: (2025)
by: Bascones, Daniel, et al.
Published: (2025)
MANTIS: A Mixed-Signal Near-Sensor Convolutional Imager SoC Using Charge-Domain 4b-Weighted 5-to-84-TOPS/W MAC Operations for Feature Extraction and Region-of-Interest Detection
by: Lefebvre, Martin, et al.
Published: (2024)
by: Lefebvre, Martin, et al.
Published: (2024)
A 2.5-nA Area-Efficient Temperature-Independent 176-/82-ppm/°C CMOS-Only Current Reference in 0.11-$μ$m Bulk and 22-nm FD-SOI
by: Lefebvre, Martin, et al.
Published: (2024)
by: Lefebvre, Martin, et al.
Published: (2024)
A nA-Range Area-Efficient Sub-100-ppm/°C Peaking Current Reference Using Forward Body Biasing in 0.11-$μ$m Bulk and 22-nm FD-SOI
by: Lefebvre, Martin, et al.
Published: (2024)
by: Lefebvre, Martin, et al.
Published: (2024)
A 1.1- / 0.9-nA Temperature-Independent 213- / 565-ppm/$^\circ$C Self-Biased CMOS-Only Current Reference in 65-nm Bulk and 22-nm FDSOI
by: Lefebvre, Martin, et al.
Published: (2023)
by: Lefebvre, Martin, et al.
Published: (2023)
HeTraX: Energy Efficient 3D Heterogeneous Manycore Architecture for Transformer Acceleration
by: Dhingra, Pratyush, et al.
Published: (2024)
by: Dhingra, Pratyush, et al.
Published: (2024)
Improving Memory Dependence Prediction with Static Analysis
by: Panayi, Luke, et al.
Published: (2024)
by: Panayi, Luke, et al.
Published: (2024)
PG-MDP: Profile-Guided Memory Dependence Prediction for Area-Constrained Cores
by: Panayi, Luke, et al.
Published: (2026)
by: Panayi, Luke, et al.
Published: (2026)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
by: Giorgi, Roberto, et al.
Published: (2026)
by: Giorgi, Roberto, et al.
Published: (2026)
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
by: Mitra, Subhasish, et al.
Published: (2025)
by: Mitra, Subhasish, et al.
Published: (2025)
basic_RV32s: An Open-Source Microarchitectural Roadmap for RISC-V RV32I
by: Kang, Hyun Woo, et al.
Published: (2025)
by: Kang, Hyun Woo, et al.
Published: (2025)
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
Data Gravity and the Energy Limits of Computation
by: Lee, Wonsuk, et al.
Published: (2026)
by: Lee, Wonsuk, et al.
Published: (2026)
Integer Representations in IEEE 754, Posit, and Takum Arithmetics
by: Hunhold, Laslo
Published: (2024)
by: Hunhold, Laslo
Published: (2024)
SoK: Where's the "up"?! A Comprehensive (bottom-up) Study on the Security of Arm Cortex-M Systems
by: Tan, Xi, et al.
Published: (2024)
by: Tan, Xi, et al.
Published: (2024)
Pinching Tactile Display: A Cloth that Changes Tactile Sensation by Electrostatic Adsorption
by: Kitagishi, Takekazu, et al.
Published: (2024)
by: Kitagishi, Takekazu, et al.
Published: (2024)
SPICEMixer - Netlist-Level Circuit Evolution
by: Uhlich, Stefan, et al.
Published: (2025)
by: Uhlich, Stefan, et al.
Published: (2025)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)
by: Kang, Hyunwoo
Published: (2026)
Minimal Neuron Circuits -- Part I: Resonators
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Minimal Neuron Circuits: Bursters
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Veryl: A New Hardware Description Language as an Altarnative to SystemVerilog
by: Hatta, Naoya, et al.
Published: (2024)
by: Hatta, Naoya, et al.
Published: (2024)
How long can you sleep? Idle Time System Inefficiencies and Opportunities
by: Antoniou, Georgia, et al.
Published: (2025)
by: Antoniou, Georgia, et al.
Published: (2025)
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
by: Shen, Fangjia, et al.
Published: (2024)
by: Shen, Fangjia, et al.
Published: (2024)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Comprehensive Formal Verification of Observational Correctness for the CHERIoT-Ibex Processor
by: Ploix, Louis-Emile, et al.
Published: (2025)
by: Ploix, Louis-Emile, et al.
Published: (2025)
Finite-Time Lyapunov Exponent Calculation on FPGA using High-Level Synthesis Tools
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Pre-Sorted Tsetlin Machine (The Genetic K-Medoid Method)
by: Morris, Jordan
Published: (2024)
by: Morris, Jordan
Published: (2024)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
by: Valpey, Benjamin, et al.
Published: (2025)
by: Valpey, Benjamin, et al.
Published: (2025)
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
ReChisel: Effective Automatic Chisel Code Generation by LLM with Reflection
by: Niu, Juxin, et al.
Published: (2025)
by: Niu, Juxin, et al.
Published: (2025)
Tekum: Balanced Ternary Tapered Precision Real Arithmetic
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
TOM: A Ternary Read-only Memory Accelerator for LLM-powered Edge Intelligence
by: Guan, Hongyi, et al.
Published: (2026)
by: Guan, Hongyi, et al.
Published: (2026)
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
by: Özyılmaz, Mustafa Mert
Published: (2026)
by: Özyılmaz, Mustafa Mert
Published: (2026)
A Flexible Instruction Set Architecture for Efficient GEMMs
by: Santana, Alexandre de Limas, et al.
Published: (2025)
by: Santana, Alexandre de Limas, et al.
Published: (2025)
Similar Items
-
An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Weight/Output Stationarity
by: Chauvaux, Nicolas, et al.
Published: (2024) -
Resource Optimized Quantum Squaring Circuit
by: Sultana, Afrin, et al.
Published: (2024) -
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
by: Goswami, Kaustav, et al.
Published: (2026) -
Chameleon: A MatMul-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
by: Blanken, Douwe den, et al.
Published: (2025) -
parti-gem5: gem5's Timing Mode Parallelised
by: Cubero-Cascante, José, et al.
Published: (2023)