Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
Fuente:
arXiv
Guardado en:
| Autores principales: | Mazzola, Sergio, Riedel, Samuel, Benini, Luca |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
por: Riedel, Samuel, et al.
Publicado: (2024)
por: Riedel, Samuel, et al.
Publicado: (2024)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
por: Mazzola, Sergio, et al.
Publicado: (2025)
por: Mazzola, Sergio, et al.
Publicado: (2025)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
por: Perotti, Matteo, et al.
Publicado: (2023)
por: Perotti, Matteo, et al.
Publicado: (2023)
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
por: Riedel, Samuel, et al.
Publicado: (2025)
por: Riedel, Samuel, et al.
Publicado: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)
por: Zhang, Yichao, et al.
Publicado: (2026)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
por: Scheffler, Paul, et al.
Publicado: (2024)
por: Scheffler, Paul, et al.
Publicado: (2024)
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
por: Wang, Bowen, et al.
Publicado: (2025)
por: Wang, Bowen, et al.
Publicado: (2025)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
por: Zhang, Yichao, et al.
Publicado: (2024)
por: Zhang, Yichao, et al.
Publicado: (2024)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
por: Leone, Lorenzo, et al.
Publicado: (2026)
por: Leone, Lorenzo, et al.
Publicado: (2026)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
por: Koenig, Cyril, et al.
Publicado: (2025)
por: Koenig, Cyril, et al.
Publicado: (2025)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
por: Mazzola, Sergio, et al.
Publicado: (2025)
por: Mazzola, Sergio, et al.
Publicado: (2025)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
por: Manoni, Simone, et al.
Publicado: (2025)
por: Manoni, Simone, et al.
Publicado: (2025)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
por: Wang, Zhao, et al.
Publicado: (2025)
por: Wang, Zhao, et al.
Publicado: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
por: Shen, Diyou, et al.
Publicado: (2025)
por: Shen, Diyou, et al.
Publicado: (2025)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
por: Rogenmoser, Michael, et al.
Publicado: (2023)
por: Rogenmoser, Michael, et al.
Publicado: (2023)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
por: Perotti, Matteo, et al.
Publicado: (2024)
por: Perotti, Matteo, et al.
Publicado: (2024)
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
por: Dhingra, Pratyush, et al.
Publicado: (2025)
por: Dhingra, Pratyush, et al.
Publicado: (2025)
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
por: Perotti, Matteo, et al.
Publicado: (2022)
por: Perotti, Matteo, et al.
Publicado: (2022)
A Direct Memory Access Controller (DMAC) for Irregular Data Transfers on RISC-V Linux Systems
por: Benz, Thomas, et al.
Publicado: (2025)
por: Benz, Thomas, et al.
Publicado: (2025)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
por: Rogenmoser, Michael, et al.
Publicado: (2023)
por: Rogenmoser, Michael, et al.
Publicado: (2023)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
por: Iff, Patrick, et al.
Publicado: (2026)
por: Iff, Patrick, et al.
Publicado: (2026)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
Ara2: Exploring Single- and Multi-Core Vector Processing with an Efficient RVV 1.0 Compliant Open-Source Processor
por: Perotti, Matteo, et al.
Publicado: (2023)
por: Perotti, Matteo, et al.
Publicado: (2023)
AraOS: Analyzing the Impact of Virtual Memory Management on Vector Unit Performance
por: Perotti, Matteo, et al.
Publicado: (2025)
por: Perotti, Matteo, et al.
Publicado: (2025)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
por: Reinwardt, Christopher, et al.
Publicado: (2025)
por: Reinwardt, Christopher, et al.
Publicado: (2025)
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads
por: Sharma, Harsh, et al.
Publicado: (2024)
por: Sharma, Harsh, et al.
Publicado: (2024)
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
por: Raja, Tejas
Publicado: (2024)
por: Raja, Tejas
Publicado: (2024)
relOBI: A Reliable Low-latency Interconnect for Tightly-Coupled On-chip Communication
por: Rogenmoser, Michael, et al.
Publicado: (2025)
por: Rogenmoser, Michael, et al.
Publicado: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
por: Verhelst, Marian, et al.
Publicado: (2025)
por: Verhelst, Marian, et al.
Publicado: (2025)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
por: Cammarata, Danilo, et al.
Publicado: (2026)
por: Cammarata, Danilo, et al.
Publicado: (2026)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
por: Perotti, Matteo, et al.
Publicado: (2024)
por: Perotti, Matteo, et al.
Publicado: (2024)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
por: Wipfli, Max, et al.
Publicado: (2026)
por: Wipfli, Max, et al.
Publicado: (2026)
Deeploy: Enabling Energy-Efficient Deployment of Small Language Models On Heterogeneous Microcontrollers
por: Scherer, Moritz, et al.
Publicado: (2024)
por: Scherer, Moritz, et al.
Publicado: (2024)
Ejemplares similares
-
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
por: Riedel, Samuel, et al.
Publicado: (2024) -
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
por: Mazzola, Sergio, et al.
Publicado: (2025) -
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
por: Perotti, Matteo, et al.
Publicado: (2023) -
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
por: Riedel, Samuel, et al.
Publicado: (2025) -
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)