TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Purayil, Navaneeth Kunhi, Shen, Diyou, Perotti, Matteo, Benini, Luca |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026)
by: Wipfli, Max, et al.
Published: (2026)
Ara2: Exploring Single- and Multi-Core Vector Processing with an Efficient RVV 1.0 Compliant Open-Source Processor
by: Perotti, Matteo, et al.
Published: (2023)
by: Perotti, Matteo, et al.
Published: (2023)
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
by: Perotti, Matteo, et al.
Published: (2022)
by: Perotti, Matteo, et al.
Published: (2022)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
by: Perotti, Matteo, et al.
Published: (2023)
by: Perotti, Matteo, et al.
Published: (2023)
AraOS: Analyzing the Impact of Virtual Memory Management on Vector Unit Performance
by: Perotti, Matteo, et al.
Published: (2025)
by: Perotti, Matteo, et al.
Published: (2025)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
by: Bertuletti, Marco, et al.
Published: (2026)
by: Bertuletti, Marco, et al.
Published: (2026)
SentryCore: A RISC-V Co-Processor System for Safe, Real-Time Control Applications
by: Rogenmoser, Michael, et al.
Published: (2024)
by: Rogenmoser, Michael, et al.
Published: (2024)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
by: Shen, Diyou, et al.
Published: (2025)
by: Shen, Diyou, et al.
Published: (2025)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
PELS: A Lightweight and Flexible Peripheral Event Linking System for Ultra-Low Power IoT Processors
by: Ottaviano, Alessandro, et al.
Published: (2023)
by: Ottaviano, Alessandro, et al.
Published: (2023)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
by: Mazzola, Sergio, et al.
Published: (2025)
by: Mazzola, Sergio, et al.
Published: (2025)
Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications
by: Cammarata, Danilo, et al.
Published: (2025)
by: Cammarata, Danilo, et al.
Published: (2025)
Quantum Hardware Roofline: Evaluating the Impact of Gate Expressivity on Quantum Processor Design
by: Kalloor, Justin, et al.
Published: (2024)
by: Kalloor, Justin, et al.
Published: (2024)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
by: Reinwardt, Christopher, et al.
Published: (2025)
by: Reinwardt, Christopher, et al.
Published: (2025)
relOBI: A Reliable Low-latency Interconnect for Tightly-Coupled On-chip Communication
by: Rogenmoser, Michael, et al.
Published: (2025)
by: Rogenmoser, Michael, et al.
Published: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
by: Bertaccini, Luca, et al.
Published: (2022)
by: Bertaccini, Luca, et al.
Published: (2022)
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
by: Riedel, Samuel, et al.
Published: (2024)
by: Riedel, Samuel, et al.
Published: (2024)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
by: Bi, Zhen, et al.
Published: (2026)
by: Bi, Zhen, et al.
Published: (2026)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
by: Ottaviano, Alessandro, et al.
Published: (2023)
by: Ottaviano, Alessandro, et al.
Published: (2023)
Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
by: Schönleber, Jannis, et al.
Published: (2023)
by: Schönleber, Jannis, et al.
Published: (2023)
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
Optimizing Offload Performance in Heterogeneous MPSoCs
by: Colagrande, Luca, et al.
Published: (2024)
by: Colagrande, Luca, et al.
Published: (2024)
DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model
by: Gerogiannis, Gerasimos, et al.
Published: (2025)
by: Gerogiannis, Gerasimos, et al.
Published: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
by: Koenig, Cyril, et al.
Published: (2025)
by: Koenig, Cyril, et al.
Published: (2025)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
by: Cammarata, Danilo, et al.
Published: (2026)
by: Cammarata, Danilo, et al.
Published: (2026)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
by: Zhang, Yichao, et al.
Published: (2026)
by: Zhang, Yichao, et al.
Published: (2026)
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
by: Bochem, Severin, et al.
Published: (2024)
by: Bochem, Severin, et al.
Published: (2024)
A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
Microarchitectural Co-Optimization for Sustained Throughput of RISC-V Multi-Lane Chaining Vector Processors
by: Wang, Weiying, et al.
Published: (2026)
by: Wang, Weiying, et al.
Published: (2026)
Similar Items
-
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025) -
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026) -
Ara2: Exploring Single- and Multi-Core Vector Processing with an Efficient RVV 1.0 Compliant Open-Source Processor
by: Perotti, Matteo, et al.
Published: (2023) -
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
by: Perotti, Matteo, et al.
Published: (2022) -
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
by: Perotti, Matteo, et al.
Published: (2024)