Dynamic Loop Fusion in High-Level Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Szafarczyk, Robert, Nabi, Syed Waqar, Vanderbauwhede, Wim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909465495732224
author Szafarczyk, Robert
Nabi, Syed Waqar
Vanderbauwhede, Wim
author_facet Szafarczyk, Robert
Nabi, Syed Waqar
Vanderbauwhede, Wim
contents Dynamic High-Level Synthesis (HLS) uses additional hardware to perform memory disambiguation at runtime, increasing loop throughput in irregular codes compared to static HLS. However, most irregular codes consist of multiple sibling loops, which currently have to be executed sequentially by all HLS tools. Static HLS performs loop fusion only on regular codes, while dynamic HLS relies on loops with dependencies to run to completion before the next loop starts. We present dynamic loop fusion for HLS, a compiler/hardware co-design approach that enables multiple loops to run in parallel, even if they contain unpredictable memory dependencies. Our only requirement is that memory addresses are monotonically non-decreasing in inner loops. We present a novel program-order schedule for HLS, inspired by polyhedral compilers, that together with our address monotonicity analysis enables dynamic memory disambiguation that does not require searching of address histories and sequential loop execution. Our evaluation shows an average speedup of 14$\times$ over static and 4$\times$ over dynamic HLS.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Loop Fusion in High-Level Synthesis
Szafarczyk, Robert
Nabi, Syed Waqar
Vanderbauwhede, Wim
Hardware Architecture
Dynamic High-Level Synthesis (HLS) uses additional hardware to perform memory disambiguation at runtime, increasing loop throughput in irregular codes compared to static HLS. However, most irregular codes consist of multiple sibling loops, which currently have to be executed sequentially by all HLS tools. Static HLS performs loop fusion only on regular codes, while dynamic HLS relies on loops with dependencies to run to completion before the next loop starts. We present dynamic loop fusion for HLS, a compiler/hardware co-design approach that enables multiple loops to run in parallel, even if they contain unpredictable memory dependencies. Our only requirement is that memory addresses are monotonically non-decreasing in inner loops. We present a novel program-order schedule for HLS, inspired by polyhedral compilers, that together with our address monotonicity analysis enables dynamic memory disambiguation that does not require searching of address histories and sequential loop execution. Our evaluation shows an average speedup of 14$\times$ over static and 4$\times$ over dynamic HLS.
title Dynamic Loop Fusion in High-Level Synthesis
topic Hardware Architecture
url https://arxiv.org/abs/2501.14631