wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hawks, Benjamin, Weitz, Jason, Demler, Dmitri, Tame-Narvaez, Karla, Plotnikov, Dennis, Rahimifar, Mohammad Mehdi, Rahali, Hamza Ezzaoui, Therrien, Audrey C., Sproule, Donovan, Khoda, Elham E, Smith, Keegan A., Marroquin, Russell, Di Guglielmo, Giuseppe, Tran, Nhan, Duarte, Javier, Loncar, Vladimir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914143368380416
author Hawks, Benjamin
Weitz, Jason
Demler, Dmitri
Tame-Narvaez, Karla
Plotnikov, Dennis
Rahimifar, Mohammad Mehdi
Rahali, Hamza Ezzaoui
Therrien, Audrey C.
Sproule, Donovan
Khoda, Elham E
Smith, Keegan A.
Marroquin, Russell
Di Guglielmo, Giuseppe
Tran, Nhan
Duarte, Javier
Loncar, Vladimir
author_facet Hawks, Benjamin
Weitz, Jason
Demler, Dmitri
Tame-Narvaez, Karla
Plotnikov, Dennis
Rahimifar, Mohammad Mehdi
Rahali, Hamza Ezzaoui
Therrien, Audrey C.
Sproule, Donovan
Khoda, Elham E
Smith, Keegan A.
Marroquin, Russell
Di Guglielmo, Giuseppe
Tran, Nhan
Duarte, Javier
Loncar, Vladimir
contents As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
Hawks, Benjamin
Weitz, Jason
Demler, Dmitri
Tame-Narvaez, Karla
Plotnikov, Dennis
Rahimifar, Mohammad Mehdi
Rahali, Hamza Ezzaoui
Therrien, Audrey C.
Sproule, Donovan
Khoda, Elham E
Smith, Keegan A.
Marroquin, Russell
Di Guglielmo, Giuseppe
Tran, Nhan
Duarte, Javier
Loncar, Vladimir
Machine Learning
Artificial Intelligence
Hardware Architecture
Instrumentation and Detectors
As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.
title wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
topic Machine Learning
Artificial Intelligence
Hardware Architecture
Instrumentation and Detectors
url https://arxiv.org/abs/2511.05615