OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Xiaoling, Antonio, Ryan, Dumoulin, Joren, Sun, Jiacong, Van Delm, Josse, Paim, Guilherme, Verhelst, Marian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
by: Antonio, Ryan Albert, et al.
Published: (2025)
by: Antonio, Ryan Albert, et al.
Published: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
by: Yi, Xiaoling, et al.
Published: (2026)
by: Yi, Xiaoling, et al.
Published: (2026)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col
by: Nayan, Md Mizanur Rahaman, et al.
Published: (2025)
by: Nayan, Md Mizanur Rahaman, et al.
Published: (2025)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
The Configuration Wall: Characterization and Elimination of Accelerator Configuration Overhead
by: Van Delm, Josse, et al.
Published: (2025)
by: Van Delm, Josse, et al.
Published: (2025)
Analog or Digital In-memory Computing? Benchmarking through Quantitative Modeling
by: Sun, Jiacong, et al.
Published: (2024)
by: Sun, Jiacong, et al.
Published: (2024)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
by: Kong, Fanchen, et al.
Published: (2025)
by: Kong, Fanchen, et al.
Published: (2025)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
by: Sarda, Giuseppe M., et al.
Published: (2025)
by: Sarda, Giuseppe M., et al.
Published: (2025)
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
by: Deng, Yunhao, et al.
Published: (2025)
by: Deng, Yunhao, et al.
Published: (2025)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
by: Shi, Man, et al.
Published: (2024)
by: Shi, Man, et al.
Published: (2024)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
by: Houshmand, Pouya, et al.
Published: (2024)
by: Houshmand, Pouya, et al.
Published: (2024)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
by: Koenig, Cyril, et al.
Published: (2025)
by: Koenig, Cyril, et al.
Published: (2025)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022)
by: Symons, Arne, et al.
Published: (2022)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
by: Mei, Linyan, et al.
Published: (2022)
by: Mei, Linyan, et al.
Published: (2022)
In-Pipeline Integration of Digital In-Memory-Computing into RISC-V Vector Architecture to Accelerate Deep Learning
by: Spagnolo, Tommaso, et al.
Published: (2026)
by: Spagnolo, Tommaso, et al.
Published: (2026)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
by: Jung, Victor J. B., et al.
Published: (2023)
by: Jung, Victor J. B., et al.
Published: (2023)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
by: Colleman, Steven, et al.
Published: (2024)
by: Colleman, Steven, et al.
Published: (2024)
X-HEEP: An Open-Source, Configurable and Extendible RISC-V Microcontroller for the Exploration of Ultra-Low-Power Edge Accelerators
by: Machetti, Simone, et al.
Published: (2024)
by: Machetti, Simone, et al.
Published: (2024)
HyperCroc: End-to-End Open-Source RISC-V MCU with a Plug-In Interface for Domain-Specific Accelerators
by: Sauter, Philippe, et al.
Published: (2026)
by: Sauter, Philippe, et al.
Published: (2026)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
by: Sarda, Giuseppe M., et al.
Published: (2024)
by: Sarda, Giuseppe M., et al.
Published: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
by: Bai, Kangbo, et al.
Published: (2025)
by: Bai, Kangbo, et al.
Published: (2025)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
by: Zniber, Alaa, et al.
Published: (2025)
by: Zniber, Alaa, et al.
Published: (2025)
MC$^2$A: Enabling Algorithm-Hardware Co-Design for Efficient Markov Chain Monte Carlo Acceleration
by: Zhao, Shirui, et al.
Published: (2025)
by: Zhao, Shirui, et al.
Published: (2025)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
by: Caon, Michele, et al.
Published: (2024)
by: Caon, Michele, et al.
Published: (2024)
Optimization of a Line Detection Algorithm for Autonomous Vehicles on a RISC-V with Accelerator
by: Belda, María José, et al.
Published: (2024)
by: Belda, María José, et al.
Published: (2024)
MARVEL: An End-to-End Framework for Generating Model-Class Aware Custom RISC-V Extensions for Lightweight AI
by: M, Ajay Kumar, et al.
Published: (2025)
by: M, Ajay Kumar, et al.
Published: (2025)
A Multi-level Compiler Backend for Accelerated Micro-kernels Targeting RISC-V ISA Extensions
by: Lopoukhine, Alexandre, et al.
Published: (2025)
by: Lopoukhine, Alexandre, et al.
Published: (2025)
CVA6S+: A Superscalar RISC-V Core with High-Throughput Memory Architecture
by: Tedeschi, Riccardo, et al.
Published: (2025)
by: Tedeschi, Riccardo, et al.
Published: (2025)
Similar Items
-
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
by: Antonio, Ryan Albert, et al.
Published: (2025) -
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
by: Yi, Xiaoling, et al.
Published: (2026) -
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025) -
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025) -
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)