Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
Fuente:
arXiv
Guardado en:
| Autores principales: | Colagrande, Luca, Leone, Lorenzo, Coco, Maximilian, Deaconeasa, Andrei, Benini, Luca |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
por: Scheffler, Paul, et al.
Publicado: (2024)
por: Scheffler, Paul, et al.
Publicado: (2024)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
por: Perotti, Matteo, et al.
Publicado: (2024)
por: Perotti, Matteo, et al.
Publicado: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
por: Manoni, Simone, et al.
Publicado: (2025)
por: Manoni, Simone, et al.
Publicado: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
Optimizing Offload Performance in Heterogeneous MPSoCs
por: Colagrande, Luca, et al.
Publicado: (2024)
por: Colagrande, Luca, et al.
Publicado: (2024)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
por: Perotti, Matteo, et al.
Publicado: (2023)
por: Perotti, Matteo, et al.
Publicado: (2023)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
por: Wipfli, Max, et al.
Publicado: (2026)
por: Wipfli, Max, et al.
Publicado: (2026)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
por: Perotti, Matteo, et al.
Publicado: (2024)
por: Perotti, Matteo, et al.
Publicado: (2024)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
por: Rogenmoser, Michael, et al.
Publicado: (2023)
por: Rogenmoser, Michael, et al.
Publicado: (2023)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
por: Koenig, Cyril, et al.
Publicado: (2025)
por: Koenig, Cyril, et al.
Publicado: (2025)
HyperCroc: End-to-End Open-Source RISC-V MCU with a Plug-In Interface for Domain-Specific Accelerators
por: Sauter, Philippe, et al.
Publicado: (2026)
por: Sauter, Philippe, et al.
Publicado: (2026)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
por: Koenig, Cyril, et al.
Publicado: (2025)
por: Koenig, Cyril, et al.
Publicado: (2025)
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
por: Rutishauser, Georg, et al.
Publicado: (2024)
por: Rutishauser, Georg, et al.
Publicado: (2024)
vCLIC: Towards Fast Interrupt Handling in Virtualized RISC-V Mixed-criticality Systems
por: Zelioli, Enrico, et al.
Publicado: (2024)
por: Zelioli, Enrico, et al.
Publicado: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
HULK-V: a Heterogeneous Ultra-low-power Linux capable RISC-V SoC
por: Valente, Luca, et al.
Publicado: (2022)
por: Valente, Luca, et al.
Publicado: (2022)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
A Direct Memory Access Controller (DMAC) for Irregular Data Transfers on RISC-V Linux Systems
por: Benz, Thomas, et al.
Publicado: (2025)
por: Benz, Thomas, et al.
Publicado: (2025)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
por: Bertaccini, Luca, et al.
Publicado: (2022)
por: Bertaccini, Luca, et al.
Publicado: (2022)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
por: Perotti, Matteo, et al.
Publicado: (2022)
por: Perotti, Matteo, et al.
Publicado: (2022)
Occamy: A 432-Core 28.1 DP-GFLOP/s/W 83% FPU Utilization Dual-Chiplet, Dual-HBM2E RISC-V-based Accelerator for Stencil and Sparse Linear Algebra Computations with 8-to-64-bit Floating-Point Support in 12nm FinFET
por: Paulin, Gianna, et al.
Publicado: (2024)
por: Paulin, Gianna, et al.
Publicado: (2024)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
por: Reinwardt, Christopher, et al.
Publicado: (2025)
por: Reinwardt, Christopher, et al.
Publicado: (2025)
Ramping Up Open-Source RISC-V Cores: Assessing the Energy Efficiency of Superscalar, Out-of-Order Execution
por: Fu, Zexin, et al.
Publicado: (2025)
por: Fu, Zexin, et al.
Publicado: (2025)
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
por: Scheffler, Paul, et al.
Publicado: (2025)
por: Scheffler, Paul, et al.
Publicado: (2025)
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
por: Jung, Victor J. B., et al.
Publicado: (2025)
por: Jung, Victor J. B., et al.
Publicado: (2025)
Occamy: A 432-Core Dual-Chiplet Dual-HBM2E 768-DP-GFLOP/s RISC-V System for 8-to-64-bit Dense and Sparse Computing in 12nm FinFET
por: Scheffler, Paul, et al.
Publicado: (2025)
por: Scheffler, Paul, et al.
Publicado: (2025)
Basilisk: An End-to-End Open-Source Linux-Capable RISC-V SoC in 130nm CMOS
por: Scheffler, Paul, et al.
Publicado: (2024)
por: Scheffler, Paul, et al.
Publicado: (2024)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
por: Mazzola, Sergio, et al.
Publicado: (2024)
por: Mazzola, Sergio, et al.
Publicado: (2024)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
por: Rogenmoser, Michael, et al.
Publicado: (2023)
por: Rogenmoser, Michael, et al.
Publicado: (2023)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
por: Wang, Run, et al.
Publicado: (2025)
por: Wang, Run, et al.
Publicado: (2025)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
por: Mazzola, Sergio, et al.
Publicado: (2025)
por: Mazzola, Sergio, et al.
Publicado: (2025)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
por: Leone, Lorenzo, et al.
Publicado: (2026)
por: Leone, Lorenzo, et al.
Publicado: (2026)
Implementing and Optimizing an Open-Source SD-card Host Controller for RISC-V SoCs
por: Vanoni, Axel, et al.
Publicado: (2026)
por: Vanoni, Axel, et al.
Publicado: (2026)
Ejemplares similares
-
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
por: Scheffler, Paul, et al.
Publicado: (2024) -
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2025) -
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2026) -
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
por: Colagrande, Luca, et al.
Publicado: (2025) -
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
por: Colagrande, Luca, et al.
Publicado: (2025)