HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Peirong, Xu, Wenchao, Wang, Haozhao, Chen, Jinyu, Shen, Xuemin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914260406239232
author Zheng, Peirong
Xu, Wenchao
Wang, Haozhao
Chen, Jinyu
Shen, Xuemin
author_facet Zheng, Peirong
Xu, Wenchao
Wang, Haozhao
Chen, Jinyu
Shen, Xuemin
contents The deployment of large language models' (LLMs) inference at the edge can facilitate prompt service responsiveness while protecting user privacy. However, it is critically challenged by the resource constraints of a single edge node. Distributed inference has emerged to aggregate and leverage computational resources across multiple devices. Yet, existing methods typically require strict synchronization, which is often infeasible due to the unreliable network conditions. In this paper, we propose HALO, a novel framework that can boost the distributed LLM inference in lossy edge network. The core idea is to enable a relaxed yet effective synchronization by strategically allocating less critical neuron groups to unstable devices, thus avoiding the excessive waiting time incurred by delayed packets. HALO introduces three key mechanisms: (1) a semantic-aware predictor to assess the significance of neuron groups prior to activation. (2) a parallel execution scheme of neuron group loading during the model inference. (3) a load-balancing scheduler that efficiently orchestrates multiple devices with heterogeneous resources. Experimental results from a Raspberry Pi cluster demonstrate that HALO achieves a 3.41x end-to-end speedup for LLaMA-series LLMs under unreliable network conditions. It maintains performance comparable to optimal conditions and significantly outperforms the state-of-the-art in various scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
Zheng, Peirong
Xu, Wenchao
Wang, Haozhao
Chen, Jinyu
Shen, Xuemin
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
The deployment of large language models' (LLMs) inference at the edge can facilitate prompt service responsiveness while protecting user privacy. However, it is critically challenged by the resource constraints of a single edge node. Distributed inference has emerged to aggregate and leverage computational resources across multiple devices. Yet, existing methods typically require strict synchronization, which is often infeasible due to the unreliable network conditions. In this paper, we propose HALO, a novel framework that can boost the distributed LLM inference in lossy edge network. The core idea is to enable a relaxed yet effective synchronization by strategically allocating less critical neuron groups to unstable devices, thus avoiding the excessive waiting time incurred by delayed packets. HALO introduces three key mechanisms: (1) a semantic-aware predictor to assess the significance of neuron groups prior to activation. (2) a parallel execution scheme of neuron group loading during the model inference. (3) a load-balancing scheduler that efficiently orchestrates multiple devices with heterogeneous resources. Experimental results from a Raspberry Pi cluster demonstrate that HALO achieves a 3.41x end-to-end speedup for LLaMA-series LLMs under unreliable network conditions. It maintains performance comparable to optimal conditions and significantly outperforms the state-of-the-art in various scenarios.
title HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
url https://arxiv.org/abs/2601.11676