Collaborative Inference for Large Models with Task Offloading and Early Exiting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Zuan, Xu, Yang, Xu, Hongli, Liao, Yunming, Yao, Zhiyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909424629579776
author Xie, Zuan
Xu, Yang
Xu, Hongli
Liao, Yunming
Yao, Zhiyuan
author_facet Xie, Zuan
Xu, Yang
Xu, Hongli
Liao, Yunming
Yao, Zhiyuan
contents In 5G smart cities, edge computing is employed to provide nearby computing services for end devices, and the large-scale models (e.g., GPT and LLaMA) can be deployed at the network edge to boost the service quality. However, due to the constraints of memory size and computing capacity, it is difficult to run these large-scale models on a single edge node. To meet the resource constraints, a large-scale model can be partitioned into multiple sub-models and deployed across multiple edge nodes. Then tasks are offloaded to the edge nodes for collaborative inference. Additionally, we incorporate the early exit mechanism to further accelerate inference. However, the heterogeneous system and dynamic environment will significantly affect the inference efficiency. To address these challenges, we theoretically analyze the coupled relationship between task offloading strategy and confidence thresholds, and develop a distributed algorithm, termed DTO-EE, based on the coupled relationship and convex optimization. DTO-EE enables each edge node to jointly optimize its offloading strategy and the confidence threshold, so as to achieve a promising trade-off between response delay and inference accuracy. The experimental results show that DTO-EE can reduce the average response delay by 21%-41% and improve the inference accuracy by 1%-4%, compared to the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08284
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Collaborative Inference for Large Models with Task Offloading and Early Exiting
Xie, Zuan
Xu, Yang
Xu, Hongli
Liao, Yunming
Yao, Zhiyuan
Distributed, Parallel, and Cluster Computing
In 5G smart cities, edge computing is employed to provide nearby computing services for end devices, and the large-scale models (e.g., GPT and LLaMA) can be deployed at the network edge to boost the service quality. However, due to the constraints of memory size and computing capacity, it is difficult to run these large-scale models on a single edge node. To meet the resource constraints, a large-scale model can be partitioned into multiple sub-models and deployed across multiple edge nodes. Then tasks are offloaded to the edge nodes for collaborative inference. Additionally, we incorporate the early exit mechanism to further accelerate inference. However, the heterogeneous system and dynamic environment will significantly affect the inference efficiency. To address these challenges, we theoretically analyze the coupled relationship between task offloading strategy and confidence thresholds, and develop a distributed algorithm, termed DTO-EE, based on the coupled relationship and convex optimization. DTO-EE enables each edge node to jointly optimize its offloading strategy and the confidence threshold, so as to achieve a promising trade-off between response delay and inference accuracy. The experimental results show that DTO-EE can reduce the average response delay by 21%-41% and improve the inference accuracy by 1%-4%, compared to the baselines.
title Collaborative Inference for Large Models with Task Offloading and Early Exiting
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2412.08284