Deep Learning Inference on Heterogeneous Mobile Processors: Potentials and Pitfalls

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Sicong, Zhou, Wentao, Zhou, Zimu, Guo, Bin, Wang, Minfan, Fang, Cheng, Lin, Zheng, Yu, Zhiwen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911864701583360
author Liu, Sicong
Zhou, Wentao
Zhou, Zimu
Guo, Bin
Wang, Minfan
Fang, Cheng
Lin, Zheng
Yu, Zhiwen
author_facet Liu, Sicong
Zhou, Wentao
Zhou, Zimu
Guo, Bin
Wang, Minfan
Fang, Cheng
Lin, Zheng
Yu, Zhiwen
contents There is a growing demand to deploy computation-intensive deep learning (DL) models on resource-constrained mobile devices for real-time intelligent applications. Equipped with a variety of processing units such as CPUs, GPUs, and NPUs, the mobile devices hold potential to accelerate DL inference via parallel execution across heterogeneous processors. Various efficient parallel methods have been explored to optimize computation distribution, achieve load balance, and minimize communication cost across processors. Yet their practical effectiveness in the dynamic and diverse real-world mobile environment is less explored. This paper presents a holistic empirical study to assess the capabilities and challenges associated with parallel DL inference on heterogeneous mobile processors. Through carefully designed experiments covering various DL models, mobile software/hardware environments, workload patterns, and resource availability, we identify limitations of existing techniques and highlight opportunities for cross-level optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2405_01851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep Learning Inference on Heterogeneous Mobile Processors: Potentials and Pitfalls
Liu, Sicong
Zhou, Wentao
Zhou, Zimu
Guo, Bin
Wang, Minfan
Fang, Cheng
Lin, Zheng
Yu, Zhiwen
Machine Learning
Artificial Intelligence
There is a growing demand to deploy computation-intensive deep learning (DL) models on resource-constrained mobile devices for real-time intelligent applications. Equipped with a variety of processing units such as CPUs, GPUs, and NPUs, the mobile devices hold potential to accelerate DL inference via parallel execution across heterogeneous processors. Various efficient parallel methods have been explored to optimize computation distribution, achieve load balance, and minimize communication cost across processors. Yet their practical effectiveness in the dynamic and diverse real-world mobile environment is less explored. This paper presents a holistic empirical study to assess the capabilities and challenges associated with parallel DL inference on heterogeneous mobile processors. Through carefully designed experiments covering various DL models, mobile software/hardware environments, workload patterns, and resource availability, we identify limitations of existing techniques and highlight opportunities for cross-level optimization.
title Deep Learning Inference on Heterogeneous Mobile Processors: Potentials and Pitfalls
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.01851