Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Taka, Endri, Gourounas, Dimitrios, Gerstlauer, Andreas, Marculescu, Diana, Arora, Aman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025)
by: Mhatre, Kaustubh, et al.
Published: (2025)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
by: Sudarshan, Chetan Choppali, et al.
Published: (2026)
by: Sudarshan, Chetan Choppali, et al.
Published: (2026)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024)
by: Xu, Han, et al.
Published: (2024)
An Energy-Efficient Artefact Detection Accelerator on FPGAs for Hyper-Spectral Satellite Imagery
by: Castelino, Cornell, et al.
Published: (2024)
by: Castelino, Cornell, et al.
Published: (2024)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)
by: Qian, Chao, et al.
Published: (2026)
LASANA: Large-Scale Surrogate Modeling for Analog Neuromorphic Architecture Exploration
by: Ho, Jason, et al.
Published: (2025)
by: Ho, Jason, et al.
Published: (2025)
MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-core Processor
by: Sui, Bingcai, et al.
Published: (2024)
by: Sui, Bingcai, et al.
Published: (2024)
Leveraging Application-Specific Knowledge for Energy-Efficient Deep Learning Accelerators on Resource-Constrained FPGAs
by: Qian, Chao
Published: (2025)
by: Qian, Chao
Published: (2025)
CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems
by: Wu, Qihang, et al.
Published: (2026)
by: Wu, Qihang, et al.
Published: (2026)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Automated Multi-Agent Workflows for RTL Design
by: Bhattaram, Amulya, et al.
Published: (2025)
by: Bhattaram, Amulya, et al.
Published: (2025)
Accelerating CRONet on AMD Versal AIE-ML Engines
by: Mhatre, Kaustubh, et al.
Published: (2026)
by: Mhatre, Kaustubh, et al.
Published: (2026)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Post-Quantum and Blockchain-Based Attestation for Trusted FPGAs in B5G Networks
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions
by: Boutros, Andrew, et al.
Published: (2024)
by: Boutros, Andrew, et al.
Published: (2024)
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification
by: Patel, Vihaan, et al.
Published: (2026)
by: Patel, Vihaan, et al.
Published: (2026)
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
by: Patras, Alexandros, et al.
Published: (2026)
by: Patras, Alexandros, et al.
Published: (2026)
Memory-efficient Sketch Acceleration for Handling Large Network Flows on FPGAs
by: Han, Zhaoyang, et al.
Published: (2025)
by: Han, Zhaoyang, et al.
Published: (2025)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
A Resource-Driven Approach for Implementing CNNs on FPGAs Using Adaptive IPs
by: Magalhães, Philippe, et al.
Published: (2025)
by: Magalhães, Philippe, et al.
Published: (2025)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
Towards Employing FPGA and ASIP Acceleration to Enable Onboard AI/ML in Space Applications
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
Wavelet Based Frequency Detection Using FPGAs
by: Hill, Caleb, et al.
Published: (2024)
by: Hill, Caleb, et al.
Published: (2024)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
by: Mao, Gang, et al.
Published: (2025)
by: Mao, Gang, et al.
Published: (2025)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
by: Hu, Jiajun, et al.
Published: (2025)
by: Hu, Jiajun, et al.
Published: (2025)
Similar Items
-
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025) -
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025) -
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
by: Taka, Endri, et al.
Published: (2025) -
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
by: Sudarshan, Chetan Choppali, et al.
Published: (2023) -
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)