Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yaodan, Zhou, Sheng, Niu, Zhisheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909586055757824
author Xu, Yaodan
Zhou, Sheng
Niu, Zhisheng
author_facet Xu, Yaodan
Zhou, Sheng
Niu, Zhisheng
contents With the growing integration of artificial intelligence in mobile applications, a substantial number of deep neural network (DNN) inference requests are generated daily by mobile devices. Serving these requests presents significant challenges due to limited device resources and strict latency requirements. Therefore, edge-device co-inference has emerged as an effective paradigm to address these issues. In this study, we focus on a scenario where multiple mobile devices offload inference tasks to an edge server equipped with a graphics processing unit (GPU). For finer control over offloading and scheduling, inference tasks are partitioned into smaller sub-tasks. Additionally, GPU batch processing is employed to boost throughput and improve energy efficiency. This work investigates the problem of minimizing total energy consumption while meeting hard latency constraints. We propose a low-complexity Joint DVFS, Offloading, and Batching strategy (J-DOB) to solve this problem. The effectiveness of the proposed algorithm is validated through extensive experiments across varying user numbers and deadline constraints. Results show that J-DOB can reduce energy consumption by up to 51.30% and 45.27% under identical and different deadlines, respectively, compared to local computing.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
Xu, Yaodan
Zhou, Sheng
Niu, Zhisheng
Distributed, Parallel, and Cluster Computing
With the growing integration of artificial intelligence in mobile applications, a substantial number of deep neural network (DNN) inference requests are generated daily by mobile devices. Serving these requests presents significant challenges due to limited device resources and strict latency requirements. Therefore, edge-device co-inference has emerged as an effective paradigm to address these issues. In this study, we focus on a scenario where multiple mobile devices offload inference tasks to an edge server equipped with a graphics processing unit (GPU). For finer control over offloading and scheduling, inference tasks are partitioned into smaller sub-tasks. Additionally, GPU batch processing is employed to boost throughput and improve energy efficiency. This work investigates the problem of minimizing total energy consumption while meeting hard latency constraints. We propose a low-complexity Joint DVFS, Offloading, and Batching strategy (J-DOB) to solve this problem. The effectiveness of the proposed algorithm is validated through extensive experiments across varying user numbers and deadline constraints. Results show that J-DOB can reduce energy consumption by up to 51.30% and 45.27% under identical and different deadlines, respectively, compared to local computing.
title Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2504.14611