A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
Fuente:
arXiv
Saved in:
| Main Authors: | Petropoulos, Anastasios, Antonakopoulos, Theodore |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
by: Bougioukou, Eleni, et al.
Published: (2025)
by: Bougioukou, Eleni, et al.
Published: (2025)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
A Scalable FPGA Architecture for Quantum Computing Simulation
by: Belfore II, Lee A.
Published: (2024)
by: Belfore II, Lee A.
Published: (2024)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
by: Parthasarathy, Rishab
Published: (2025)
by: Parthasarathy, Rishab
Published: (2025)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
by: Lee, Dongjae, et al.
Published: (2025)
by: Lee, Dongjae, et al.
Published: (2025)
FPGA-based Hyrbid Memory Emulation System
by: Wen, Fei, et al.
Published: (2020)
by: Wen, Fei, et al.
Published: (2020)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
by: Ferry, Corentin, et al.
Published: (2024)
by: Ferry, Corentin, et al.
Published: (2024)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
by: Caon, Michele, et al.
Published: (2024)
by: Caon, Michele, et al.
Published: (2024)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
by: Ko, Younghoon, et al.
Published: (2026)
by: Ko, Younghoon, et al.
Published: (2026)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
An FPGA Compiler for On-the-Fly Adaptive CNN Deployment and Reconfiguration
by: Mazouz, Alaa, et al.
Published: (2025)
by: Mazouz, Alaa, et al.
Published: (2025)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-core Processor
by: Sui, Bingcai, et al.
Published: (2024)
by: Sui, Bingcai, et al.
Published: (2024)
VolTune: A Fine-Grained Runtime Voltage Control Architecture for FPGA Systems
by: Ahmed, Akram Ben, et al.
Published: (2026)
by: Ahmed, Akram Ben, et al.
Published: (2026)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
by: Pun, Junius, et al.
Published: (2025)
by: Pun, Junius, et al.
Published: (2025)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
Adaptive Hybrid FFT: A Novel Pipeline and Memory-Based Architecture for Radix-$2^k$ FFT in Large Size Processing
by: Zhao, Fangyu, et al.
Published: (2025)
by: Zhao, Fangyu, et al.
Published: (2025)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
by: Gerlinghoff, Daniel, et al.
Published: (2024)
by: Gerlinghoff, Daniel, et al.
Published: (2024)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025)
by: Mhatre, Kaustubh, et al.
Published: (2025)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
by: Gao, Yizhao, et al.
Published: (2024)
by: Gao, Yizhao, et al.
Published: (2024)
A Low-Dissipation and Scalable GEMM Accelerator with Silicon Nitride Photonics
by: Karempudi, Venkata Sai Praneeth, et al.
Published: (2024)
by: Karempudi, Venkata Sai Praneeth, et al.
Published: (2024)
LaZagna: An Open-Source Framework for Flexible 3D FPGA Architectural Exploration
by: Youssef, Ismael, et al.
Published: (2025)
by: Youssef, Ismael, et al.
Published: (2025)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
by: Sun, Jialin, et al.
Published: (2025)
by: Sun, Jialin, et al.
Published: (2025)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
by: Chen, Yiqi, et al.
Published: (2025)
by: Chen, Yiqi, et al.
Published: (2025)
A Hybrid Residue Floating Numerical Architecture with Formal Error Bounds for High Throughput FPGA Computation
by: Darvishi, Mostafa
Published: (2026)
by: Darvishi, Mostafa
Published: (2026)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Development of High-Performance DSP Algorithms on the European Rad-Hard NG-ULTRA SoC FPGA
by: Leon, Vasileios, et al.
Published: (2024)
by: Leon, Vasileios, et al.
Published: (2024)
A Scalable Resource Management Layer for FPGA SoCs in 6G Radio Units
by: Bartzoudis, Nikolaos, et al.
Published: (2025)
by: Bartzoudis, Nikolaos, et al.
Published: (2025)
FPPS: An FPGA-Based Point Cloud Processing System
by: Zhou, Xiaofeng, et al.
Published: (2026)
by: Zhou, Xiaofeng, et al.
Published: (2026)
Design and Implementation of BNN-Based Object Detection on FPGA
by: Zhao, Xuyu, et al.
Published: (2026)
by: Zhao, Xuyu, et al.
Published: (2026)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
by: Cammarata, Danilo, et al.
Published: (2026)
by: Cammarata, Danilo, et al.
Published: (2026)
SCREME: A Scalable Framework for Resilient Memory Design
by: Li, Fan, et al.
Published: (2025)
by: Li, Fan, et al.
Published: (2025)
SPAC: Automating FPGA-based Network Switches with Protocol Adaptive Customization
by: Li, Guoyu, et al.
Published: (2026)
by: Li, Guoyu, et al.
Published: (2026)
Similar Items
-
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025) -
Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
by: Bougioukou, Eleni, et al.
Published: (2025) -
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025) -
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025) -
A Scalable FPGA Architecture for Quantum Computing Simulation
by: Belfore II, Lee A.
Published: (2024)