PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Xiang, Xu, Zikang, Qi, Qi, Wang, Jingyu, Sun, Haifeng, Liao, Jianxin, Guo, Song
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909148298346496
author Yang, Xiang
Xu, Zikang
Qi, Qi
Wang, Jingyu
Sun, Haifeng
Liao, Jianxin
Guo, Song
author_facet Yang, Xiang
Xu, Zikang
Qi, Qi
Wang, Jingyu
Sun, Haifeng
Liao, Jianxin
Guo, Song
contents Distributing the inference of convolutional neural network (CNN) to multiple mobile devices has been studied in recent years to achieve real-time inference without losing accuracy. However, how to map CNN to devices remains a challenge. On the one hand, scheduling the workload of state-of-the-art CNNs with multiple devices is NP-Hard because the structures of CNNs are directed acyclic graphs (DAG) rather than simple chains. On the other hand, distributing the inference workload suffers from expensive communication and unbalanced computation due to the wireless environment and heterogeneous devices. This paper presents PICO, a pipeline cooperation framework to accelerate the inference of versatile CNNs on diverse mobile devices. At its core, PICO features: (1) a generic graph partition algorithm that considers the characteristics of any given CNN and orchestrates it into a list of model pieces with suitable granularity, and (2) a many-to-many mapping algorithm that produces the best pipeline configuration for heterogeneous devices. In our experiment with 2 ~ 8 Raspberry-Pi devices, the throughput can be improved by 1.8 ~ 6.8x under different CPU frequencies.
format Preprint
id arxiv_https___arxiv_org_abs_2206_08662
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
Yang, Xiang
Xu, Zikang
Qi, Qi
Wang, Jingyu
Sun, Haifeng
Liao, Jianxin
Guo, Song
Distributed, Parallel, and Cluster Computing
Distributing the inference of convolutional neural network (CNN) to multiple mobile devices has been studied in recent years to achieve real-time inference without losing accuracy. However, how to map CNN to devices remains a challenge. On the one hand, scheduling the workload of state-of-the-art CNNs with multiple devices is NP-Hard because the structures of CNNs are directed acyclic graphs (DAG) rather than simple chains. On the other hand, distributing the inference workload suffers from expensive communication and unbalanced computation due to the wireless environment and heterogeneous devices. This paper presents PICO, a pipeline cooperation framework to accelerate the inference of versatile CNNs on diverse mobile devices. At its core, PICO features: (1) a generic graph partition algorithm that considers the characteristics of any given CNN and orchestrates it into a list of model pieces with suitable granularity, and (2) a many-to-many mapping algorithm that produces the best pipeline configuration for heterogeneous devices. In our experiment with 2 ~ 8 Raspberry-Pi devices, the throughput can be improved by 1.8 ~ 6.8x under different CPU frequencies.
title PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2206.08662