MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Que, Zhiqiang, Coutinho, Jose G. F., Guo, Ce, Fan, Hongxiang, Luk, Wayne
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917261574406144
author Que, Zhiqiang
Coutinho, Jose G. F.
Guo, Ce
Fan, Hongxiang
Luk, Wayne
author_facet Que, Zhiqiang
Coutinho, Jose G. F.
Guo, Ce
Fan, Hongxiang
Luk, Wayne
contents This paper presents a unified framework for codifying and automating optimization strategies to efficiently deploy deep neural networks (DNNs) on resource-constrained hardware, such as FPGAs, while maintaining high performance, accuracy, and resource efficiency. Deploying DNNs on such platforms involves addressing the significant challenge of balancing performance, resource usage (e.g., DSPs and LUTs), and inference accuracy, which often requires extensive manual effort and domain expertise. Our novel approach addresses two core key issues: (i)~encoding custom optimization strategies and (ii)~enabling cross-stage optimization search. In particular, our proposed framework seamlessly integrates programmatic DNN optimization techniques with high-level synthesis (HLS)-based metaprogramming, leveraging advanced design space exploration (DSE) strategies like Bayesian optimization to automate both top-down and bottom-up design flows. Hence, we reduce the need for manual intervention and domain expertise. In addition, the framework introduces customizable optimization, transformation, and control blocks to enhance DNN accelerator performance and resource efficiency. Experimental results demonstrate up to a 92\% DSP and 89\% LUT usage reduction for select networks, while preserving accuracy, along with a 15.6-fold reduction in optimization time compared to grid search. These results highlight the potential for automating the generation of resource-efficient DNN accelerator designs with minimum effort.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05850
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
Que, Zhiqiang
Coutinho, Jose G. F.
Guo, Ce
Fan, Hongxiang
Luk, Wayne
Hardware Architecture
Machine Learning
This paper presents a unified framework for codifying and automating optimization strategies to efficiently deploy deep neural networks (DNNs) on resource-constrained hardware, such as FPGAs, while maintaining high performance, accuracy, and resource efficiency. Deploying DNNs on such platforms involves addressing the significant challenge of balancing performance, resource usage (e.g., DSPs and LUTs), and inference accuracy, which often requires extensive manual effort and domain expertise. Our novel approach addresses two core key issues: (i)~encoding custom optimization strategies and (ii)~enabling cross-stage optimization search. In particular, our proposed framework seamlessly integrates programmatic DNN optimization techniques with high-level synthesis (HLS)-based metaprogramming, leveraging advanced design space exploration (DSE) strategies like Bayesian optimization to automate both top-down and bottom-up design flows. Hence, we reduce the need for manual intervention and domain expertise. In addition, the framework introduces customizable optimization, transformation, and control blocks to enhance DNN accelerator performance and resource efficiency. Experimental results demonstrate up to a 92\% DSP and 89\% LUT usage reduction for select networks, while preserving accuracy, along with a 15.6-fold reduction in optimization time compared to grid search. These results highlight the potential for automating the generation of resource-efficient DNN accelerator designs with minimum effort.
title MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2502.05850