Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kreß, Fabian, Annabi, El Mahdi El, Hotfilter, Tim, Hoefer, Julian, Harbaum, Tanja, Becker, Juergen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929537236860928
author Kreß, Fabian
Annabi, El Mahdi El
Hotfilter, Tim
Hoefer, Julian
Harbaum, Tanja
Becker, Juergen
author_facet Kreß, Fabian
Annabi, El Mahdi El
Hotfilter, Tim
Hoefer, Julian
Harbaum, Tanja
Becker, Juergen
contents Distributed systems can be found in various applications, e.g., in robotics or autonomous driving, to achieve higher flexibility and robustness. Thereby, data flow centric applications such as Deep Neural Network (DNN) inference benefit from partitioning the workload over multiple compute nodes in terms of performance and energy-efficiency. However, mapping large models on distributed embedded systems is a complex task, due to low latency and high throughput requirements combined with strict energy and memory constraints. In this paper, we present a novel approach for hardware-aware layer scheduling of DNN inference in distributed embedded systems. Therefore, our proposed framework uses a graph-based algorithm to automatically find beneficial partitioning points in a given DNN. Each of these is evaluated based on several essential system metrics such as accuracy and memory utilization, while considering the respective system constraints. We demonstrate our approach in terms of the impact of inference partitioning on various performance metrics of six different DNNs. As an example, we can achieve a 47.5 % throughput increase for EfficientNet-B0 inference partitioned onto two platforms while observing high energy-efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
Kreß, Fabian
Annabi, El Mahdi El
Hotfilter, Tim
Hoefer, Julian
Harbaum, Tanja
Becker, Juergen
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Distributed systems can be found in various applications, e.g., in robotics or autonomous driving, to achieve higher flexibility and robustness. Thereby, data flow centric applications such as Deep Neural Network (DNN) inference benefit from partitioning the workload over multiple compute nodes in terms of performance and energy-efficiency. However, mapping large models on distributed embedded systems is a complex task, due to low latency and high throughput requirements combined with strict energy and memory constraints. In this paper, we present a novel approach for hardware-aware layer scheduling of DNN inference in distributed embedded systems. Therefore, our proposed framework uses a graph-based algorithm to automatically find beneficial partitioning points in a given DNN. Each of these is evaluated based on several essential system metrics such as accuracy and memory utilization, while considering the respective system constraints. We demonstrate our approach in terms of the impact of inference partitioning on various performance metrics of six different DNNs. As an example, we can achieve a 47.5 % throughput increase for EfficientNet-B0 inference partitioned onto two platforms while observing high energy-efficiency.
title Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2406.19913