Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Symons, Arne, Mei, Linyan, Colleman, Steven, Houshmand, Pouya, Karl, Sebastian, Verhelst, Marian
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909827510304768
author Symons, Arne
Mei, Linyan
Colleman, Steven
Houshmand, Pouya
Karl, Sebastian
Verhelst, Marian
author_facet Symons, Arne
Mei, Linyan
Colleman, Steven
Houshmand, Pouya
Karl, Sebastian
Verhelst, Marian
contents As the landscape of deep neural networks evolves, heterogeneous dataflow accelerators, in the form of multi-core architectures or chiplet-based designs, promise more flexibility and higher inference performance through scalability. So far, these systems exploit the increased parallelism by coarsely mapping a single layer at a time across cores, which incurs frequent costly off-chip memory accesses, or by pipelining batches of inputs, which falls short in meeting the demands of latency-critical applications. To alleviate these bottlenecks, this work explores a new fine-grain mapping paradigm, referred to as layer fusion, on heterogeneous dataflow accelerators through a novel design space exploration framework called Stream. Stream captures a wide variety of heterogeneous dataflow architectures and mapping granularities, and implements a memory and communication-aware latency and energy analysis validated with three distinct state-of-the-art hardware implementations. As such, it facilitates a holistic exploration of architecture and mapping, by strategically allocating the workload through constraint optimization. The findings demonstrate that the integration of layer fusion with heterogeneous dataflow accelerators yields up to 2.2x lower energy-delay product in inference efficiency, addressing both energy consumption and latency concerns. The framework is available open-source at: https://github.com/kuleuven-micas/stream.
format Preprint
id arxiv_https___arxiv_org_abs_2212_10612
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
Symons, Arne
Mei, Linyan
Colleman, Steven
Houshmand, Pouya
Karl, Sebastian
Verhelst, Marian
Hardware Architecture
As the landscape of deep neural networks evolves, heterogeneous dataflow accelerators, in the form of multi-core architectures or chiplet-based designs, promise more flexibility and higher inference performance through scalability. So far, these systems exploit the increased parallelism by coarsely mapping a single layer at a time across cores, which incurs frequent costly off-chip memory accesses, or by pipelining batches of inputs, which falls short in meeting the demands of latency-critical applications. To alleviate these bottlenecks, this work explores a new fine-grain mapping paradigm, referred to as layer fusion, on heterogeneous dataflow accelerators through a novel design space exploration framework called Stream. Stream captures a wide variety of heterogeneous dataflow architectures and mapping granularities, and implements a memory and communication-aware latency and energy analysis validated with three distinct state-of-the-art hardware implementations. As such, it facilitates a holistic exploration of architecture and mapping, by strategically allocating the workload through constraint optimization. The findings demonstrate that the integration of layer fusion with heterogeneous dataflow accelerators yields up to 2.2x lower energy-delay product in inference efficiency, addressing both energy consumption and latency concerns. The framework is available open-source at: https://github.com/kuleuven-micas/stream.
title Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
topic Hardware Architecture
url https://arxiv.org/abs/2212.10612