Design Rules for Extreme-Edge Scientific Computing on AI Engines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Zhenghua, Abarajithan, G, Danopoulos, Dimitrios, Weng, Olivia, Restuccia, Francesco, Kastner, Ryan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908982958882816
author Ma, Zhenghua
Abarajithan, G
Danopoulos, Dimitrios
Weng, Olivia
Restuccia, Francesco
Kastner, Ryan
author_facet Ma, Zhenghua
Abarajithan, G
Danopoulos, Dimitrios
Weng, Olivia
Restuccia, Francesco
Kastner, Ryan
contents Extreme-edge scientific applications use machine learning models to analyze sensor data and make real-time decisions. Their stringent latency and throughput requirements demand small batch sizes and require that model weights remain fully on-chip. Spatial dataflow implementations are common for extreme-edge applications. Spatial dataflow works well for small networks, but it fails to scale to larger models due to inherent resource scaling limitations. AI Engines on modern FPGA SoCs offer a promising alternative with high compute density and additional on-chip memory. However, the architecture, programming model, and performance-scaling behavior of AI Engines differ fundamentally from those of the programmable logic, making direct comparison non-trivial and the benefits of using AI Engines unclear. This work addresses how and when extreme-edge scientific neural networks should be implemented on AI Engines versus programmable logic. We provide systematic architectural characterization and micro-benchmarking and introduce a latency-adjusted resource equivalence (LARE) metric that identifies when AI Engine implementations outperform programmable logic designs. We further propose spatial and API-level dataflow optimizations tailored to low-latency scientific inference. Finally, we demonstrate the successful deployment of end-to-end neural networks on AI Engines that cannot fit on programmable logic when using the hlsml toolchain.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19106
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Design Rules for Extreme-Edge Scientific Computing on AI Engines
Ma, Zhenghua
Abarajithan, G
Danopoulos, Dimitrios
Weng, Olivia
Restuccia, Francesco
Kastner, Ryan
Hardware Architecture
Artificial Intelligence
Machine Learning
Extreme-edge scientific applications use machine learning models to analyze sensor data and make real-time decisions. Their stringent latency and throughput requirements demand small batch sizes and require that model weights remain fully on-chip. Spatial dataflow implementations are common for extreme-edge applications. Spatial dataflow works well for small networks, but it fails to scale to larger models due to inherent resource scaling limitations. AI Engines on modern FPGA SoCs offer a promising alternative with high compute density and additional on-chip memory. However, the architecture, programming model, and performance-scaling behavior of AI Engines differ fundamentally from those of the programmable logic, making direct comparison non-trivial and the benefits of using AI Engines unclear. This work addresses how and when extreme-edge scientific neural networks should be implemented on AI Engines versus programmable logic. We provide systematic architectural characterization and micro-benchmarking and introduce a latency-adjusted resource equivalence (LARE) metric that identifies when AI Engine implementations outperform programmable logic designs. We further propose spatial and API-level dataflow optimizations tailored to low-latency scientific inference. Finally, we demonstrate the successful deployment of end-to-end neural networks on AI Engines that cannot fit on programmable logic when using the hlsml toolchain.
title Design Rules for Extreme-Edge Scientific Computing on AI Engines
topic Hardware Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.19106