Fast and Practical Strassen's Matrix Multiplication using FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmad, Afzal, Du, Linfeng, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A WASM-Subset Stack Architecture for Low-cost FPGAs using Open-Source EDA Flows
by: Chakrabarti, Aradhya
Published: (2025)
by: Chakrabarti, Aradhya
Published: (2025)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025)
by: Taji, Hossein, et al.
Published: (2025)
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
by: Zhang, Yuhan, et al.
Published: (2026)
by: Zhang, Yuhan, et al.
Published: (2026)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
by: Prabhakar, Raghu, et al.
Published: (2024)
by: Prabhakar, Raghu, et al.
Published: (2024)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
FREESS: An Educational Simulator of a RISC-V-Inspired Superscalar Processor Based on Tomasulo's Algorithm
by: Giorgi, Roberto
Published: (2025)
by: Giorgi, Roberto
Published: (2025)
SISA: A Scale-In Systolic Array for GEMM Acceleration
by: Altamura, Luigi, et al.
Published: (2026)
by: Altamura, Luigi, et al.
Published: (2026)
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
by: Lee, Ming-Yen, et al.
Published: (2025)
by: Lee, Ming-Yen, et al.
Published: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
by: Suryadevara, Pranav
Published: (2025)
by: Suryadevara, Pranav
Published: (2025)
A Compilation Framework for Quantum Circuits with Mid-Circuit Measurement Error Awareness
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
Non-interfering On-line and In-field SoC Testing
by: Strauch, Tobias
Published: (2024)
by: Strauch, Tobias
Published: (2024)
Research on LLM Acceleration Using the High-Performance RISC-V Processor "Xiangshan" (Nanhu Version) Based on the Open-Source Matrix Instruction Set Extension (Vector Dot Product)
by: Chen, Xu-Hao, et al.
Published: (2024)
by: Chen, Xu-Hao, et al.
Published: (2024)
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
by: Das, Tamoghno, et al.
Published: (2025)
by: Das, Tamoghno, et al.
Published: (2025)
Multi-diseases detection with memristive system on chip
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026)
by: Kyriazis, Agamemnon, et al.
Published: (2026)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
by: Ferreira, João Dinis, et al.
Published: (2021)
by: Ferreira, João Dinis, et al.
Published: (2021)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026)
by: Villarrubia, Jorge, et al.
Published: (2026)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
by: Das, Debabrata, et al.
Published: (2026)
by: Das, Debabrata, et al.
Published: (2026)
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
by: Koeplinger, David, et al.
Published: (2024)
by: Koeplinger, David, et al.
Published: (2024)
FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
by: Parameshwara, Arya, et al.
Published: (2025)
by: Parameshwara, Arya, et al.
Published: (2025)
Time Domain Near Memory Computing Engine
by: Antal, Sarthak, et al.
Published: (2026)
by: Antal, Sarthak, et al.
Published: (2026)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
by: Siracusa, Marco, et al.
Published: (2025)
by: Siracusa, Marco, et al.
Published: (2025)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation
by: Abouagour, Faris
Published: (2026)
by: Abouagour, Faris
Published: (2026)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
Leveraging FPGAs for Homomorphic Matrix-Vector Multiplication in Oblivious Message Retrieval
by: Bosworth, Grant, et al.
Published: (2025)
by: Bosworth, Grant, et al.
Published: (2025)
DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
Improved Prefetching Techniques for Linked Data Structures
by: Maruszewski, Nikola Vuk
Published: (2025)
by: Maruszewski, Nikola Vuk
Published: (2025)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
by: Campos, Nelson, et al.
Published: (2024)
by: Campos, Nelson, et al.
Published: (2024)
A Scalable Architecture for Efficient Multi-bit Fully Homomorphic Encryption
by: Ma, Jiaao, et al.
Published: (2025)
by: Ma, Jiaao, et al.
Published: (2025)
RedMulE-FT: A Reconfigurable Fault-Tolerant Matrix Multiplication Engine
by: Wiese, Philip, et al.
Published: (2025)
by: Wiese, Philip, et al.
Published: (2025)
C for a tiny system
by: Krause, Philipp Klaus, et al.
Published: (2020)
by: Krause, Philipp Klaus, et al.
Published: (2020)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
AMC: Access to Miss Correlation Prefetcher for Evolving Graph Analytics
by: Singh, Abhishek, et al.
Published: (2024)
by: Singh, Abhishek, et al.
Published: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
FPPS: An FPGA-Based Point Cloud Processing System
by: Zhou, Xiaofeng, et al.
Published: (2026)
by: Zhou, Xiaofeng, et al.
Published: (2026)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
Similar Items
-
A WASM-Subset Stack Architecture for Low-cost FPGAs using Open-Source EDA Flows
by: Chakrabarti, Aradhya
Published: (2025) -
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025) -
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025) -
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
by: Zhang, Yuhan, et al.
Published: (2026) -
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)