Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
Fuente:
arXiv
Saved in:
| Main Authors: | Schönleber, Jannis, Cavigelli, Lukas, Perotti, Matteo, Benini, Luca, Andri, Renzo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ara2: Exploring Single- and Multi-Core Vector Processing with an Efficient RVV 1.0 Compliant Open-Source Processor
by: Perotti, Matteo, et al.
Published: (2023)
by: Perotti, Matteo, et al.
Published: (2023)
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
by: Perotti, Matteo, et al.
Published: (2022)
by: Perotti, Matteo, et al.
Published: (2022)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
by: Perotti, Matteo, et al.
Published: (2023)
by: Perotti, Matteo, et al.
Published: (2023)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Explicit Sign-Magnitude Encoders Enable Power-Efficient Multipliers
by: Arnold, Felix, et al.
Published: (2025)
by: Arnold, Felix, et al.
Published: (2025)
The Art of Beating the Odds with Predictor-Guided Random Design Space Exploration
by: Arnold, Felix, et al.
Published: (2025)
by: Arnold, Felix, et al.
Published: (2025)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units
by: Bouvier, Maxence, et al.
Published: (2025)
by: Bouvier, Maxence, et al.
Published: (2025)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
Energy Efficient Exact and Approximate Systolic Array Architecture for Matrix Multiplication
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
hARMS: A Hardware Acceleration Architecture for Real-Time Event-Based Optical Flow
by: Stumpp, Daniel C., et al.
Published: (2021)
by: Stumpp, Daniel C., et al.
Published: (2021)
AraOS: Analyzing the Impact of Virtual Memory Management on Vector Unit Performance
by: Perotti, Matteo, et al.
Published: (2025)
by: Perotti, Matteo, et al.
Published: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
SentryCore: A RISC-V Co-Processor System for Safe, Real-Time Control Applications
by: Rogenmoser, Michael, et al.
Published: (2024)
by: Rogenmoser, Michael, et al.
Published: (2024)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications
by: Cammarata, Danilo, et al.
Published: (2025)
by: Cammarata, Danilo, et al.
Published: (2025)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
by: Fu, Yonggan, et al.
Published: (2023)
by: Fu, Yonggan, et al.
Published: (2023)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
by: Scheffler, Paul, et al.
Published: (2024)
by: Scheffler, Paul, et al.
Published: (2024)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
by: Wang, Run, et al.
Published: (2026)
by: Wang, Run, et al.
Published: (2026)
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
by: Li, Shuaiting, et al.
Published: (2024)
by: Li, Shuaiting, et al.
Published: (2024)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026)
by: Wipfli, Max, et al.
Published: (2026)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
by: Bertaccini, Luca, et al.
Published: (2022)
by: Bertaccini, Luca, et al.
Published: (2022)
SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
by: Zhang, Yipu, et al.
Published: (2025)
by: Zhang, Yipu, et al.
Published: (2025)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
by: Carrigg, Kieran, et al.
Published: (2026)
by: Carrigg, Kieran, et al.
Published: (2026)
Design and Analysis of Approximate Hardware Accelerators for VVC Intra Angular Prediction
by: de Fraga, Lucas M. Leipnitz, et al.
Published: (2025)
by: de Fraga, Lucas M. Leipnitz, et al.
Published: (2025)
GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping
by: Eudine, Julien, et al.
Published: (2026)
by: Eudine, Julien, et al.
Published: (2026)
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
by: Zniber, Alaa, et al.
Published: (2025)
by: Zniber, Alaa, et al.
Published: (2025)
HyperCroc: End-to-End Open-Source RISC-V MCU with a Plug-In Interface for Domain-Specific Accelerators
by: Sauter, Philippe, et al.
Published: (2026)
by: Sauter, Philippe, et al.
Published: (2026)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
by: Koenig, Cyril, et al.
Published: (2025)
by: Koenig, Cyril, et al.
Published: (2025)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
by: Koenig, Cyril, et al.
Published: (2025)
by: Koenig, Cyril, et al.
Published: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
by: Su, Yuchao, et al.
Published: (2025)
by: Su, Yuchao, et al.
Published: (2025)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
by: Manoni, Simone, et al.
Published: (2025)
by: Manoni, Simone, et al.
Published: (2025)
Similar Items
-
Ara2: Exploring Single- and Multi-Core Vector Processing with an Efficient RVV 1.0 Compliant Open-Source Processor
by: Perotti, Matteo, et al.
Published: (2023) -
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
by: Perotti, Matteo, et al.
Published: (2022) -
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
by: Zhang, Chi, et al.
Published: (2026) -
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
by: Perotti, Matteo, et al.
Published: (2024) -
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
by: Perotti, Matteo, et al.
Published: (2023)