DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiayi, Da Lu, Ang, Zeng, Zhichen, Li, Ang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Capstone: Power-Capped Pipelining for Coarse-Grained Reconfigurable Array Compilers
by: Yarzada, Sabrina, et al.
Published: (2026)
by: Yarzada, Sabrina, et al.
Published: (2026)
Mapping code on Coarse Grained Reconfigurable Arrays using a SAT solver
by: Tirelli, Cristian, et al.
Published: (2025)
by: Tirelli, Cristian, et al.
Published: (2025)
SAT-MapIt: A SAT-based Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures
by: Tirelli, Cristian, et al.
Published: (2025)
by: Tirelli, Cristian, et al.
Published: (2025)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Banked Memories for Soft SIMT Processors
by: Langhammer, Martin, et al.
Published: (2025)
by: Langhammer, Martin, et al.
Published: (2025)
HeLEx: A Heterogeneous Layout Explorer for Spatial Elastic Coarse-Grained Reconfigurable Arrays
by: Du, Alan Jia Bao, et al.
Published: (2025)
by: Du, Alan Jia Bao, et al.
Published: (2025)
A 950 MHz SIMT Soft Processor
by: Langhammer, Martin, et al.
Published: (2025)
by: Langhammer, Martin, et al.
Published: (2025)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures
by: Li, Shangkun, et al.
Published: (2026)
by: Li, Shangkun, et al.
Published: (2026)
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
by: Langarita, Rubén, et al.
Published: (2025)
by: Langarita, Rubén, et al.
Published: (2025)
Guac: Energy-Aware and SSA-Based Generation of Coarse-Grained Merged Accelerators from LLVM-IR
by: Brumar, Iulian, et al.
Published: (2024)
by: Brumar, Iulian, et al.
Published: (2024)
Reconfigurable Digital RRAM Logic Enables In-Situ Pruning and Learning for Edge AI
by: Wang, Songqi, et al.
Published: (2025)
by: Wang, Songqi, et al.
Published: (2025)
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
by: Walter, Dominik, et al.
Published: (2025)
by: Walter, Dominik, et al.
Published: (2025)
MASIM: An Efficient Multi-Array Scheduler for In-Memory SIMD Computation
by: Qian, Xingyue, et al.
Published: (2024)
by: Qian, Xingyue, et al.
Published: (2024)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design
by: Wang, Kainan, et al.
Published: (2025)
by: Wang, Kainan, et al.
Published: (2025)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution Redundancy
by: Tang, Jiaping, et al.
Published: (2025)
by: Tang, Jiaping, et al.
Published: (2025)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
by: Huang, Sixiao, et al.
Published: (2025)
by: Huang, Sixiao, et al.
Published: (2025)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
by: Kim, Jae-Young, et al.
Published: (2025)
by: Kim, Jae-Young, et al.
Published: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
by: Han, Meng, et al.
Published: (2023)
by: Han, Meng, et al.
Published: (2023)
Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator
by: Ando, Takuto, et al.
Published: (2025)
by: Ando, Takuto, et al.
Published: (2025)
AssertMiner: Module-Level Spec Generation and Assertion Mining using Static Analysis Guided LLMs
by: Lyu, Hongqin, et al.
Published: (2025)
by: Lyu, Hongqin, et al.
Published: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
by: Li, Tenglong, et al.
Published: (2024)
by: Li, Tenglong, et al.
Published: (2024)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
Reconfigurable Stream Network Architecture
by: Wang, Chengyue, et al.
Published: (2024)
by: Wang, Chengyue, et al.
Published: (2024)
Energy-Efficient QoS-Aware Scheduling for S-NUCA Many-Cores
by: Wasala, Sudam M., et al.
Published: (2025)
by: Wasala, Sudam M., et al.
Published: (2025)
On Reducing the Execution Latency of Superconducting Quantum Processors via Quantum Job Scheduling
by: Wu, Wenjie, et al.
Published: (2024)
by: Wu, Wenjie, et al.
Published: (2024)
VersaQ-3D: A Reconfigurable Accelerator Enabling Feed-Forward and Generalizable 3D Reconstruction via Versatile Quantization
by: Zhang, Yipu, et al.
Published: (2026)
by: Zhang, Yipu, et al.
Published: (2026)
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
by: Chen, Bo-Yu, et al.
Published: (2025)
by: Chen, Bo-Yu, et al.
Published: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
by: Su, Yuchao, et al.
Published: (2025)
by: Su, Yuchao, et al.
Published: (2025)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
by: Jang, Yongjoo, et al.
Published: (2025)
by: Jang, Yongjoo, et al.
Published: (2025)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
by: Min, Kyeongpil, et al.
Published: (2026)
by: Min, Kyeongpil, et al.
Published: (2026)
Trinity: A General Purpose FHE Accelerator
by: Deng, Xianglong, et al.
Published: (2024)
by: Deng, Xianglong, et al.
Published: (2024)
A Novel Cost-Effective MIMO Architecture with Ray Antenna Array for Enhanced Wireless Communication Performance
by: Dong, Zhenjun, et al.
Published: (2025)
by: Dong, Zhenjun, et al.
Published: (2025)
Similar Items
-
Capstone: Power-Capped Pipelining for Coarse-Grained Reconfigurable Array Compilers
by: Yarzada, Sabrina, et al.
Published: (2026) -
Mapping code on Coarse Grained Reconfigurable Arrays using a SAT solver
by: Tirelli, Cristian, et al.
Published: (2025) -
SAT-MapIt: A SAT-based Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures
by: Tirelli, Cristian, et al.
Published: (2025) -
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025) -
Banked Memories for Soft SIMT Processors
by: Langhammer, Martin, et al.
Published: (2025)