Saved in:
Bibliographic Details
Main Authors: Gao, Yunquan, Zhang, Zhiguo, Donta, Praveen Kumar, Dehury, Chinmaya Kumar, Wang, Xiujun, Niyato, Dusit, Zhang, Qiyang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.21109
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916663587241984
author Gao, Yunquan
Zhang, Zhiguo
Donta, Praveen Kumar
Dehury, Chinmaya Kumar
Wang, Xiujun
Niyato, Dusit
Zhang, Qiyang
author_facet Gao, Yunquan
Zhang, Zhiguo
Donta, Praveen Kumar
Dehury, Chinmaya Kumar
Wang, Xiujun
Niyato, Dusit
Zhang, Qiyang
contents Deep Neural Networks (DNNs) are increasingly deployed across diverse industries, driving demand for mobile device support. However, existing mobile inference frameworks often rely on a single processor per model, limiting hardware utilization and causing suboptimal performance and energy efficiency. Expanding DNN accessibility on mobile platforms requires adaptive, resource-efficient solutions to meet rising computational needs without compromising functionality. Parallel inference of multiple DNNs on heterogeneous processors remains challenging. Some works partition DNN operations into subgraphs for parallel execution across processors, but these often create excessive subgraphs based only on hardware compatibility, increasing scheduling complexity and memory overhead. To address this, we propose an Advanced Multi-DNN Model Scheduling (ADMS) strategy for optimizing multi-DNN inference on mobile heterogeneous processors. ADMS constructs an optimal subgraph partitioning strategy offline, balancing hardware operation support and scheduling granularity, and uses a processor-state-aware algorithm to dynamically adjust workloads based on real-time conditions. This ensures efficient workload distribution and maximizes processor utilization. Experiments show ADMS reduces multi-DNN inference latency by 4.04 times compared to vanilla frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21109
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
Gao, Yunquan
Zhang, Zhiguo
Donta, Praveen Kumar
Dehury, Chinmaya Kumar
Wang, Xiujun
Niyato, Dusit
Zhang, Qiyang
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
68T07, 68W40
I.2.6; C.1.4; D.4.8
Deep Neural Networks (DNNs) are increasingly deployed across diverse industries, driving demand for mobile device support. However, existing mobile inference frameworks often rely on a single processor per model, limiting hardware utilization and causing suboptimal performance and energy efficiency. Expanding DNN accessibility on mobile platforms requires adaptive, resource-efficient solutions to meet rising computational needs without compromising functionality. Parallel inference of multiple DNNs on heterogeneous processors remains challenging. Some works partition DNN operations into subgraphs for parallel execution across processors, but these often create excessive subgraphs based only on hardware compatibility, increasing scheduling complexity and memory overhead. To address this, we propose an Advanced Multi-DNN Model Scheduling (ADMS) strategy for optimizing multi-DNN inference on mobile heterogeneous processors. ADMS constructs an optimal subgraph partitioning strategy offline, balancing hardware operation support and scheduling granularity, and uses a processor-state-aware algorithm to dynamically adjust workloads based on real-time conditions. This ensures efficient workload distribution and maximizes processor utilization. Experiments show ADMS reduces multi-DNN inference latency by 4.04 times compared to vanilla frameworks.
title Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
68T07, 68W40
I.2.6; C.1.4; D.4.8
url https://arxiv.org/abs/2503.21109