How to keep pushing ML accelerator performance? Know your rooflines!
Fuente:
arXiv
Guardado en:
| Autores principales: | Verhelst, Marian, Benini, Luca, Verma, Naveen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
por: Houshmand, Pouya, et al.
Publicado: (2024)
por: Houshmand, Pouya, et al.
Publicado: (2024)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
por: Geens, Robin, et al.
Publicado: (2025)
por: Geens, Robin, et al.
Publicado: (2025)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
por: Colleman, Steven, et al.
Publicado: (2024)
por: Colleman, Steven, et al.
Publicado: (2024)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
por: Jung, Victor J. B., et al.
Publicado: (2023)
por: Jung, Victor J. B., et al.
Publicado: (2023)
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
por: Geens, Robin, et al.
Publicado: (2026)
por: Geens, Robin, et al.
Publicado: (2026)
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
por: Geens, Robin, et al.
Publicado: (2025)
por: Geens, Robin, et al.
Publicado: (2025)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
por: Cammarata, Danilo, et al.
Publicado: (2026)
por: Cammarata, Danilo, et al.
Publicado: (2026)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
por: Dumoulin, Joren, et al.
Publicado: (2025)
por: Dumoulin, Joren, et al.
Publicado: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
por: Geens, Robin, et al.
Publicado: (2026)
por: Geens, Robin, et al.
Publicado: (2026)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
por: Rogenmoser, Michael, et al.
Publicado: (2023)
por: Rogenmoser, Michael, et al.
Publicado: (2023)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
por: Sarda, Giuseppe M., et al.
Publicado: (2025)
por: Sarda, Giuseppe M., et al.
Publicado: (2025)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
por: Sarda, Giuseppe M., et al.
Publicado: (2024)
por: Sarda, Giuseppe M., et al.
Publicado: (2024)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
por: Yi, Xiaoling, et al.
Publicado: (2025)
por: Yi, Xiaoling, et al.
Publicado: (2025)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
por: Symons, Arne, et al.
Publicado: (2022)
por: Symons, Arne, et al.
Publicado: (2022)
relOBI: A Reliable Low-latency Interconnect for Tightly-Coupled On-chip Communication
por: Rogenmoser, Michael, et al.
Publicado: (2025)
por: Rogenmoser, Michael, et al.
Publicado: (2025)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
por: Koenig, Cyril, et al.
Publicado: (2025)
por: Koenig, Cyril, et al.
Publicado: (2025)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
por: Mazzola, Sergio, et al.
Publicado: (2024)
por: Mazzola, Sergio, et al.
Publicado: (2024)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
por: Cuyckens, Stef, et al.
Publicado: (2025)
por: Cuyckens, Stef, et al.
Publicado: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
por: Yi, Xiaoling, et al.
Publicado: (2026)
por: Yi, Xiaoling, et al.
Publicado: (2026)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
por: Cuyckens, Stef, et al.
Publicado: (2025)
por: Cuyckens, Stef, et al.
Publicado: (2025)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
TOP: Towards Open & Predictable Heterogeneous SoCs
por: Valente, Luca, et al.
Publicado: (2024)
por: Valente, Luca, et al.
Publicado: (2024)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
por: Bertaccini, Luca, et al.
Publicado: (2022)
por: Bertaccini, Luca, et al.
Publicado: (2022)
A Direct Memory Access Controller (DMAC) for Irregular Data Transfers on RISC-V Linux Systems
por: Benz, Thomas, et al.
Publicado: (2025)
por: Benz, Thomas, et al.
Publicado: (2025)
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
por: Riedel, Samuel, et al.
Publicado: (2025)
por: Riedel, Samuel, et al.
Publicado: (2025)
HyperCroc: End-to-End Open-Source RISC-V MCU with a Plug-In Interface for Domain-Specific Accelerators
por: Sauter, Philippe, et al.
Publicado: (2026)
por: Sauter, Philippe, et al.
Publicado: (2026)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
por: Perotti, Matteo, et al.
Publicado: (2023)
por: Perotti, Matteo, et al.
Publicado: (2023)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
por: Koenig, Cyril, et al.
Publicado: (2025)
por: Koenig, Cyril, et al.
Publicado: (2025)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
por: Manoni, Simone, et al.
Publicado: (2025)
por: Manoni, Simone, et al.
Publicado: (2025)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
por: Scheffler, Paul, et al.
Publicado: (2024)
por: Scheffler, Paul, et al.
Publicado: (2024)
Analog or Digital In-memory Computing? Benchmarking through Quantitative Modeling
por: Sun, Jiacong, et al.
Publicado: (2024)
por: Sun, Jiacong, et al.
Publicado: (2024)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
Ejemplares similares
-
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
por: Houshmand, Pouya, et al.
Publicado: (2024) -
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
por: Geens, Robin, et al.
Publicado: (2025) -
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
por: Colleman, Steven, et al.
Publicado: (2024) -
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
por: Jung, Victor J. B., et al.
Publicado: (2023) -
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
por: Geens, Robin, et al.
Publicado: (2026)