Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Bin, Banerjee, Ayan, Gupta, Sandeep
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909946710327296
author Xu, Bin
Banerjee, Ayan
Gupta, Sandeep
author_facet Xu, Bin
Banerjee, Ayan
Gupta, Sandeep
contents Model Recovery (MR) is a core primitive for physical AI and real-time digital twins, but GPUs often execute MR inefficiently due to iterative dependencies, kernel-launch overheads, underutilized memory bandwidth, and high data-movement latency. We present MERINDA, an FPGA-accelerated MR framework that restructures computation as a streaming dataflow pipeline. MERINDA exploits on-chip locality through BRAM tiling, fixed-point kernels, and the concurrent use of LUT fabric and carry-chain adders to expose fine-grained spatial parallelism while minimizing off-chip traffic. This hardware-aware formulation removes synchronization bottlenecks and sustains high throughput across the iterative updates in MR. On representative MR workloads, MERINDA delivers up to 6.3x fewer cycles than an FPGA-based LTC baseline, enabling real-time performance for time-critical physical systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
Xu, Bin
Banerjee, Ayan
Gupta, Sandeep
Hardware Architecture
Machine Learning
Model Recovery (MR) is a core primitive for physical AI and real-time digital twins, but GPUs often execute MR inefficiently due to iterative dependencies, kernel-launch overheads, underutilized memory bandwidth, and high data-movement latency. We present MERINDA, an FPGA-accelerated MR framework that restructures computation as a streaming dataflow pipeline. MERINDA exploits on-chip locality through BRAM tiling, fixed-point kernels, and the concurrent use of LUT fabric and carry-chain adders to expose fine-grained spatial parallelism while minimizing off-chip traffic. This hardware-aware formulation removes synchronization bottlenecks and sustains high throughput across the iterative updates in MR. On representative MR workloads, MERINDA delivers up to 6.3x fewer cycles than an FPGA-based LTC baseline, enabling real-time performance for time-critical physical systems.
title Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2512.06113