Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jindong, Li, Tenglong, Shen, Guobin, Zhao, Dongcheng, Zhang, Qian, Zeng, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025)
von: Li, Jindong, et al.
Veröffentlicht: (2025)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025)
von: Li, Jindong, et al.
Veröffentlicht: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
FPGA & VPU Co-Processing in Space Applications: Development and Testing with DSP/AI Benchmarks
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
von: Taka, Endri, et al.
Veröffentlicht: (2025)
von: Taka, Endri, et al.
Veröffentlicht: (2025)
Development of High-Performance DSP Algorithms on the European Rad-Hard NG-ULTRA SoC FPGA
von: Leon, Vasileios, et al.
Veröffentlicht: (2024)
von: Leon, Vasileios, et al.
Veröffentlicht: (2024)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
von: Raja, Tejas
Veröffentlicht: (2024)
von: Raja, Tejas
Veröffentlicht: (2024)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
Design and Implementation of BNN-Based Object Detection on FPGA
von: Zhao, Xuyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xuyu, et al.
Veröffentlicht: (2026)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
von: Wu, Qizhe, et al.
Veröffentlicht: (2024)
von: Wu, Qizhe, et al.
Veröffentlicht: (2024)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
A DSP shared is a DSP earned: HLS Task-Level Multi-Pumping for High-Performance Low-Resource Designs
von: Brignone, Giovanni, et al.
Veröffentlicht: (2023)
von: Brignone, Giovanni, et al.
Veröffentlicht: (2023)
Holistic Optimization Framework for FPGA Accelerators
von: Pouget, Stéphane, et al.
Veröffentlicht: (2025)
von: Pouget, Stéphane, et al.
Veröffentlicht: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
von: Ding, Hong, et al.
Veröffentlicht: (2025)
von: Ding, Hong, et al.
Veröffentlicht: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
Error Checking for Sparse Systolic Tensor Arrays
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
Structural Mutation Based Differential Testing for FPGA Logic Synthesis Compilers
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
FPPS: An FPGA-Based Point Cloud Processing System
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2026)
Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
von: Huang, Mingqiang, et al.
Veröffentlicht: (2024)
von: Huang, Mingqiang, et al.
Veröffentlicht: (2024)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
von: He, Zicheng, et al.
Veröffentlicht: (2026)
von: He, Zicheng, et al.
Veröffentlicht: (2026)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
von: Ferry, Corentin, et al.
Veröffentlicht: (2024)
von: Ferry, Corentin, et al.
Veröffentlicht: (2024)
Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
Graphitron: A Domain Specific Language for FPGA-based Graph Processing Accelerator Generation
von: Zhang, Xinmiao, et al.
Veröffentlicht: (2024)
von: Zhang, Xinmiao, et al.
Veröffentlicht: (2024)
FASE: FPGA-Assisted Syscall Emulation for Rapid End-to-End Processor Performance Validation
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
Online Learning Extreme Learning Machine with Low-Complexity Predictive Plasticity Rule and FPGA Implementation
von: Zang, Zhenya, et al.
Veröffentlicht: (2025)
von: Zang, Zhenya, et al.
Veröffentlicht: (2025)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
von: Yin, Chenyang, et al.
Veröffentlicht: (2025)
von: Yin, Chenyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
von: Li, Tenglong, et al.
Veröffentlicht: (2026) -
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025) -
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2025) -
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025) -
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)