Intra-DP: A High Performance Collaborative Inference System for Mobile Edge Computing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Zekai, Guan, Xiuxian, Lin, Zheng, Fang, Zihan, Cai, Xiangming, Chen, Zhe, Liu, Fangming, Cui, Heming, Xiong, Jie, Ni, Wei, Yuen, Chau
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914051849715712
author Sun, Zekai
Guan, Xiuxian
Lin, Zheng
Fang, Zihan
Cai, Xiangming
Chen, Zhe
Liu, Fangming
Cui, Heming
Xiong, Jie
Ni, Wei
Yuen, Chau
author_facet Sun, Zekai
Guan, Xiuxian
Lin, Zheng
Fang, Zihan
Cai, Xiangming
Chen, Zhe
Liu, Fangming
Cui, Heming
Xiong, Jie
Ni, Wei
Yuen, Chau
contents Deploying deep neural networks (DNNs) on resource-constrained mobile devices presents significant challenges, particularly in achieving real-time performance while simultaneously coping with limited computational resources and battery life. While Mobile Edge Computing (MEC) offers collaborative inference with GPU servers as a promising solution, existing approaches primarily rely on layer-wise model partitioning and undergo significant transmission bottlenecks caused by the sequential execution of DNN operations. To address this challenge, we present Intra-DP, a high-performance collaborative inference system optimized for DNN inference on MEC. Intra DP employs a novel parallel computing technique based on local operators (i.e., operators whose minimum unit input is not the entire input tensor, such as the convolution kernel). By decomposing their computations (operations) into several independent sub-operations and overlapping the computation and transmission of different sub-operations through parallel execution, Intra-DP mitigates transmission bottlenecks in MEC, achieving fast and energy-efficient inference. The evaluation demonstrates that Intra-DP reduces per-inference latency by up to 50% and energy consumption by up to 75% compared to state-of-the-art baselines, without sacrificing accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Intra-DP: A High Performance Collaborative Inference System for Mobile Edge Computing
Sun, Zekai
Guan, Xiuxian
Lin, Zheng
Fang, Zihan
Cai, Xiangming
Chen, Zhe
Liu, Fangming
Cui, Heming
Xiong, Jie
Ni, Wei
Yuen, Chau
Networking and Internet Architecture
Artificial Intelligence
Machine Learning
Deploying deep neural networks (DNNs) on resource-constrained mobile devices presents significant challenges, particularly in achieving real-time performance while simultaneously coping with limited computational resources and battery life. While Mobile Edge Computing (MEC) offers collaborative inference with GPU servers as a promising solution, existing approaches primarily rely on layer-wise model partitioning and undergo significant transmission bottlenecks caused by the sequential execution of DNN operations. To address this challenge, we present Intra-DP, a high-performance collaborative inference system optimized for DNN inference on MEC. Intra DP employs a novel parallel computing technique based on local operators (i.e., operators whose minimum unit input is not the entire input tensor, such as the convolution kernel). By decomposing their computations (operations) into several independent sub-operations and overlapping the computation and transmission of different sub-operations through parallel execution, Intra-DP mitigates transmission bottlenecks in MEC, achieving fast and energy-efficient inference. The evaluation demonstrates that Intra-DP reduces per-inference latency by up to 50% and energy consumption by up to 75% compared to state-of-the-art baselines, without sacrificing accuracy.
title Intra-DP: A High Performance Collaborative Inference System for Mobile Edge Computing
topic Networking and Internet Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.05829