DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Xin, Fan, Xin, Wang, Zengshi, Han, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
FREESS: An Educational Simulator of a RISC-V-Inspired Superscalar Processor Based on Tomasulo's Algorithm
by: Giorgi, Roberto
Published: (2025)
by: Giorgi, Roberto
Published: (2025)
SISA: A Scale-In Systolic Array for GEMM Acceleration
by: Altamura, Luigi, et al.
Published: (2026)
by: Altamura, Luigi, et al.
Published: (2026)
Static Communication Analysis for Hardware Design
by: Rosendahl, Mads, et al.
Published: (2025)
by: Rosendahl, Mads, et al.
Published: (2025)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
by: Siracusa, Marco, et al.
Published: (2025)
by: Siracusa, Marco, et al.
Published: (2025)
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026)
by: Kyriazis, Agamemnon, et al.
Published: (2026)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025)
by: Behrendt, Emet, et al.
Published: (2025)
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
by: Das, Tamoghno, et al.
Published: (2025)
by: Das, Tamoghno, et al.
Published: (2025)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
by: Parameshwara, Arya, et al.
Published: (2025)
by: Parameshwara, Arya, et al.
Published: (2025)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
by: Vahdatniya, Parmida, et al.
Published: (2025)
by: Vahdatniya, Parmida, et al.
Published: (2025)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
by: Ahmad, Afzal, et al.
Published: (2024)
by: Ahmad, Afzal, et al.
Published: (2024)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)
by: Kang, Hyunwoo
Published: (2026)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
Design and implementation of a synchronous Hardware Performance Monitor for a RISC-V space-oriented processor
by: Arribas, Miguel Jiménez, et al.
Published: (2024)
by: Arribas, Miguel Jiménez, et al.
Published: (2024)
ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
by: Durvasula, Sankeerth, et al.
Published: (2024)
by: Durvasula, Sankeerth, et al.
Published: (2024)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025)
by: Jeong, Shinnung, et al.
Published: (2025)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
by: Bertaccini, Luca, et al.
Published: (2022)
by: Bertaccini, Luca, et al.
Published: (2022)
IPU: Flexible Hardware Introspection Units
by: McDougall, Ian, et al.
Published: (2023)
by: McDougall, Ian, et al.
Published: (2023)
AMC: Access to Miss Correlation Prefetcher for Evolving Graph Analytics
by: Singh, Abhishek, et al.
Published: (2024)
by: Singh, Abhishek, et al.
Published: (2024)
AGON: Automated Design Framework for Customizing Processors from ISA Documents
by: Li, Chongxiao, et al.
Published: (2024)
by: Li, Chongxiao, et al.
Published: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
by: Khadem, Alireza, et al.
Published: (2025)
by: Khadem, Alireza, et al.
Published: (2025)
On the Impact of ISA Extension on Energy Consumption of I-Cache in Extensible Processors
by: Behboudi, Noushin, et al.
Published: (2024)
by: Behboudi, Noushin, et al.
Published: (2024)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
How long can you sleep? Idle Time System Inefficiencies and Opportunities
by: Antoniou, Georgia, et al.
Published: (2025)
by: Antoniou, Georgia, et al.
Published: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
by: Zhang, Yuhan, et al.
Published: (2026)
by: Zhang, Yuhan, et al.
Published: (2026)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
Similar Items
-
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025) -
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025) -
FREESS: An Educational Simulator of a RISC-V-Inspired Superscalar Processor Based on Tomasulo's Algorithm
by: Giorgi, Roberto
Published: (2025) -
SISA: A Scale-In Systolic Array for GEMM Acceleration
by: Altamura, Luigi, et al.
Published: (2026) -
Static Communication Analysis for Hardware Design
by: Rosendahl, Mads, et al.
Published: (2025)