A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Petropoulos, Anastasios, Antonakopoulos, Theodore
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918157488226304
author Petropoulos, Anastasios
Antonakopoulos, Theodore
author_facet Petropoulos, Anastasios
Antonakopoulos, Theodore
contents Deep neural network (DNN) inference relies increasingly on specialized hardware for high computational efficiency. This work introduces a field-programmable gate array (FPGA)-based dynamically configurable accelerator featuring systolic arrays, high-bandwidth memory, and UltraRAMs. We present two processing unit (PU) configurations with different computing capabilities using the same interfaces and peripheral blocks. By instantiating multiple PUs and employing a heuristic weight transfer schedule, the architecture achieves notable throughput efficiency over prior works. Moreover, we outline how the architecture can be extended to emulate analog in-memory computing (AIMC) devices to aid next-generation heterogeneous AIMC chip designs and investigate device-level noise behavior. Overall, this brief presents a versatile DNN inference acceleration architecture adaptable to various models and future FPGA designs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
Petropoulos, Anastasios
Antonakopoulos, Theodore
Hardware Architecture
Deep neural network (DNN) inference relies increasingly on specialized hardware for high computational efficiency. This work introduces a field-programmable gate array (FPGA)-based dynamically configurable accelerator featuring systolic arrays, high-bandwidth memory, and UltraRAMs. We present two processing unit (PU) configurations with different computing capabilities using the same interfaces and peripheral blocks. By instantiating multiple PUs and employing a heuristic weight transfer schedule, the architecture achieves notable throughput efficiency over prior works. Moreover, we outline how the architecture can be extended to emulate analog in-memory computing (AIMC) devices to aid next-generation heterogeneous AIMC chip designs and investigate device-level noise behavior. Overall, this brief presents a versatile DNN inference acceleration architecture adaptable to various models and future FPGA designs.
title A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
topic Hardware Architecture
url https://arxiv.org/abs/2510.08137