Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dagli, Ismet, Belviranli, Mehmet
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917583684370432
author Dagli, Ismet
Belviranli, Mehmet
author_facet Dagli, Ismet
Belviranli, Mehmet
contents Two distinguishing features of state-of-the-art mobile and autonomous systems are 1) there are often multiple workloads, mainly deep neural network (DNN) inference, running concurrently and continuously; and 2) they operate on shared memory system-on-chips (SoC) that embed heterogeneous accelerators tailored for specific operations. State-of-the-art lacks efficient performance and resource management techniques necessary to either maximize total system throughput or minimize end-to-end workload latency. In this work, we propose HaX-CoNN, a novel scheme that characterizes and maps layers in concurrently executing DNN inference workloads to a diverse set of accelerators within a SoC. Our scheme uniquely takes per-layer execution characteristics, shared memory (SM) contention, and inter-accelerator transitions into account to find optimal schedules. We evaluate HaX-CoNN on NVIDIA Orin, NVIDIA Xavier, and Qualcomm Snapdragon 865 SoCs. Our experimental results indicate that HaX-CoNN minimizes memory contention by up to 45% and can improve latency and total throughput by up to 32% and 29%, respectively, compared to the state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2308_05869
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
Dagli, Ismet
Belviranli, Mehmet
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Performance
Two distinguishing features of state-of-the-art mobile and autonomous systems are 1) there are often multiple workloads, mainly deep neural network (DNN) inference, running concurrently and continuously; and 2) they operate on shared memory system-on-chips (SoC) that embed heterogeneous accelerators tailored for specific operations. State-of-the-art lacks efficient performance and resource management techniques necessary to either maximize total system throughput or minimize end-to-end workload latency. In this work, we propose HaX-CoNN, a novel scheme that characterizes and maps layers in concurrently executing DNN inference workloads to a diverse set of accelerators within a SoC. Our scheme uniquely takes per-layer execution characteristics, shared memory (SM) contention, and inter-accelerator transitions into account to find optimal schedules. We evaluate HaX-CoNN on NVIDIA Orin, NVIDIA Xavier, and Qualcomm Snapdragon 865 SoCs. Our experimental results indicate that HaX-CoNN minimizes memory contention by up to 45% and can improve latency and total throughput by up to 32% and 29%, respectively, compared to the state-of-the-art approaches.
title Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Performance
url https://arxiv.org/abs/2308.05869