HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Peirong, Xu, Wenchao, Wang, Haozhao, Chen, Jinyu, Shen, Xuemin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
by: Nguyen, Chanh, et al.
Published: (2026)
by: Nguyen, Chanh, et al.
Published: (2026)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Edge Graph Intelligence: Reciprocally Empowering Edge Networks with Graph Intelligence
by: Zeng, Liekang, et al.
Published: (2024)
by: Zeng, Liekang, et al.
Published: (2024)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025)
by: Luo, Haoxiang, et al.
Published: (2025)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)
by: Yang, Zheming, et al.
Published: (2024)
Design and Optimization of Hierarchical Gradient Coding for Distributed Learning at Edge Devices
by: Tang, Weiheng, et al.
Published: (2024)
by: Tang, Weiheng, et al.
Published: (2024)
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
by: Ahmadpanah, Seyed Hossein
Published: (2025)
by: Ahmadpanah, Seyed Hossein
Published: (2025)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
by: Ye, Shengyuan, et al.
Published: (2025)
by: Ye, Shengyuan, et al.
Published: (2025)
Network Anomaly Detection in Distributed Edge Computing Infrastructure
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
Contention-Aware Microservice Deployment in Collaborative Mobile Edge Networks
by: Ge, Xinlei, et al.
Published: (2024)
by: Ge, Xinlei, et al.
Published: (2024)
Early-Exit meets Model-Distributed Inference at Edge Networks
by: Colocrese, Marco, et al.
Published: (2024)
by: Colocrese, Marco, et al.
Published: (2024)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
by: Liyanage, Mohan, et al.
Published: (2026)
by: Liyanage, Mohan, et al.
Published: (2026)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
by: Zhang, Haijun, et al.
Published: (2025)
by: Zhang, Haijun, et al.
Published: (2025)
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
by: Wang, Zhanwei, et al.
Published: (2026)
by: Wang, Zhanwei, et al.
Published: (2026)
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
by: Yang, Jin, et al.
Published: (2025)
by: Yang, Jin, et al.
Published: (2025)
SANSee: A Physical-layer Semantic-aware Networking Framework for Distributed Wireless Sensing
by: Zhu, Huixiang, et al.
Published: (2024)
by: Zhu, Huixiang, et al.
Published: (2024)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Diffusion Models on the Edge: Challenges, Optimizations, and Applications
by: Zheng, Dongqi
Published: (2025)
by: Zheng, Dongqi
Published: (2025)
PeerSync: Accelerating Containerized Service Delivery at the Network Edge
by: Deng, Yinuo, et al.
Published: (2025)
by: Deng, Yinuo, et al.
Published: (2025)
Topology-aware Microservice Architecture in Edge Networks: Deployment Optimization and Implementation
by: Chen, Yuang, et al.
Published: (2025)
by: Chen, Yuang, et al.
Published: (2025)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
by: Han, Mingqi, et al.
Published: (2026)
by: Han, Mingqi, et al.
Published: (2026)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
Placing Timely Refreshing Services at the Network Edge
by: Li, Xishuo, et al.
Published: (2024)
by: Li, Xishuo, et al.
Published: (2024)
A Multi-Layered Distributed Computing Framework for Enhanced Edge Computing
by: Ma, Ke, et al.
Published: (2024)
by: Ma, Ke, et al.
Published: (2024)
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
by: Sun, Tingyang, et al.
Published: (2025)
by: Sun, Tingyang, et al.
Published: (2025)
Towards Timely Video Analytics Services at the Network Edge
by: Li, Xishuo, et al.
Published: (2024)
by: Li, Xishuo, et al.
Published: (2024)
GENIO: Synergizing Edge Computing with Optical Network Infrastructures
by: Cesarano, Carmine, et al.
Published: (2025)
by: Cesarano, Carmine, et al.
Published: (2025)
Surviving the Edge: Federated Learning under Networking and Resource Constraints
by: Mwanje, Mike, et al.
Published: (2026)
by: Mwanje, Mike, et al.
Published: (2026)
Dynamic Edge Server Selection in Time-Varying Environments: A Reliability-Aware Predictive Approach
by: Burbano, Jaime Sebastian, et al.
Published: (2025)
by: Burbano, Jaime Sebastian, et al.
Published: (2025)
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
by: Sahni, Yuvraj, et al.
Published: (2024)
by: Sahni, Yuvraj, et al.
Published: (2024)
Decentralized Network Topology Design for Task Offloading in Mobile Edge Computing
by: Ma, Ke, et al.
Published: (2024)
by: Ma, Ke, et al.
Published: (2024)
Deduplicator: When Computation Reuse Meets Load Balancing at the Network Edge
by: Azad, Md Washik Al, et al.
Published: (2024)
by: Azad, Md Washik Al, et al.
Published: (2024)
Towards Edge General Intelligence via Large Language Models: Opportunities and Challenges
by: Chen, Handi, et al.
Published: (2024)
by: Chen, Handi, et al.
Published: (2024)
Distributed Simulation for Digital Twins of Large-Scale Real-World DiffServ-Based Networks
by: Huang, Zhuoyao, et al.
Published: (2024)
by: Huang, Zhuoyao, et al.
Published: (2024)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
by: Jang, SiYoung, et al.
Published: (2025)
by: Jang, SiYoung, et al.
Published: (2025)
The AutoSPADA Platform: User-Friendly Edge Computing for Distributed Learning and Data Analytics in Connected Vehicles
by: Nilsson, Adrian, et al.
Published: (2023)
by: Nilsson, Adrian, et al.
Published: (2023)
ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference Modeling
by: Krishnamachari, Bhaskar, et al.
Published: (2026)
by: Krishnamachari, Bhaskar, et al.
Published: (2026)
Dynamic DAG-Application Scheduling for Multi-Tier Edge Computing in Heterogeneous Networks
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
D-LoRa: a Distributed Parameter Adaptation Scheme for LoRa Network
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
by: Sivtsov, Danil, et al.
Published: (2025)
by: Sivtsov, Danil, et al.
Published: (2025)
Similar Items
-
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
by: Nguyen, Chanh, et al.
Published: (2026) -
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025) -
Edge Graph Intelligence: Reciprocally Empowering Edge Networks with Graph Intelligence
by: Zeng, Liekang, et al.
Published: (2024) -
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025) -
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)