Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tayal, Mumuksh, Simmhan, Yogesh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916650953998336
author Tayal, Mumuksh
Simmhan, Yogesh
author_facet Tayal, Mumuksh
Simmhan, Yogesh
contents Edge devices like Nvidia Jetson platforms now offer several on-board accelerators -- including GPU CUDA cores, Tensor Cores, and Deep Learning Accelerators (DLA) -- which can be concurrently exploited to boost deep neural network (DNN) inferencing. In this paper, we extend previous work by evaluating the performance impacts of running multiple instances of the ResNet50 model concurrently across these heterogeneous components. We detail the effects of varying batch sizes and hardware combinations on throughput and latency. Our expanded analysis highlights not only the benefits of combining CUDA and Tensor Cores, but also the performance degradation from resource contention when integrating DLAs. These findings, together with insights on precision constraints and workload allocation challenges, motivate further exploration of intelligent scheduling mechanisms to optimize resource utilization on edge platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09546
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
Tayal, Mumuksh
Simmhan, Yogesh
Distributed, Parallel, and Cluster Computing
Edge devices like Nvidia Jetson platforms now offer several on-board accelerators -- including GPU CUDA cores, Tensor Cores, and Deep Learning Accelerators (DLA) -- which can be concurrently exploited to boost deep neural network (DNN) inferencing. In this paper, we extend previous work by evaluating the performance impacts of running multiple instances of the ResNet50 model concurrently across these heterogeneous components. We detail the effects of varying batch sizes and hardware combinations on throughput and latency. Our expanded analysis highlights not only the benefits of combining CUDA and Tensor Cores, but also the performance degradation from resource contention when integrating DLAs. These findings, together with insights on precision constraints and workload allocation challenges, motivate further exploration of intelligent scheduling mechanisms to optimize resource utilization on edge platforms.
title Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.09546