Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhixiong, Zhu, Bingjie, Wang, Jiangzhou, Shin, Hyundong, Nallanathan, Arumugam, Niyato, Dusit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
by: Qiu, Houming, et al.
Published: (2025)
by: Qiu, Houming, et al.
Published: (2025)
Lightweight Federated Learning over Wireless Edge Networks
by: Hou, Xiangwang, et al.
Published: (2025)
by: Hou, Xiangwang, et al.
Published: (2025)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
Approximated Coded Computing: Towards Fast, Private and Secure Distributed Machine Learning
by: Qiu, Houming, et al.
Published: (2024)
by: Qiu, Houming, et al.
Published: (2024)
Federated Unlearning in Edge Networks: A Survey of Fundamentals, Challenges, Practical Applications and Future Directions
by: Ng, Jer Shyuan, et al.
Published: (2026)
by: Ng, Jer Shyuan, et al.
Published: (2026)
Edge Association Strategies for Synthetic Data Empowered Hierarchical Federated Learning with Non-IID Data
by: Ng, Jer Shyuan, et al.
Published: (2025)
by: Ng, Jer Shyuan, et al.
Published: (2025)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
by: Hakim, Sheikh Azizul, et al.
Published: (2025)
by: Hakim, Sheikh Azizul, et al.
Published: (2025)
Energy Efficient Federated Learning with Hyperdimensional Computing over Wireless Communication Networks
by: Ding, Yahao, et al.
Published: (2026)
by: Ding, Yahao, et al.
Published: (2026)
Unleashing the Power of Tree-of-Thoughts for Edge-Enabled AIGC Service Provisioning
by: Liu, Zhang, et al.
Published: (2026)
by: Liu, Zhang, et al.
Published: (2026)
Data Heterogeneity-Aware Client Selection for Federated Learning in Wireless Networks
by: Yang, Yanbing, et al.
Published: (2025)
by: Yang, Yanbing, et al.
Published: (2025)
Federated Customization of Large Models: Approaches, Experiments, and Insights
by: Ye, Yuchuan, et al.
Published: (2026)
by: Ye, Yuchuan, et al.
Published: (2026)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025)
by: Luo, Haoxiang, et al.
Published: (2025)
A Review on Edge Large Language Models: Design, Execution, and Applications
by: Zheng, Yue, et al.
Published: (2024)
by: Zheng, Yue, et al.
Published: (2024)
Knowledge-driven Reasoning for Mobile Agentic AI: Concepts, Approaches, and Directions
by: Liu, Guangyuan, et al.
Published: (2026)
by: Liu, Guangyuan, et al.
Published: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
by: Chen, Fahao, et al.
Published: (2025)
by: Chen, Fahao, et al.
Published: (2025)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
Large Language Model Partitioning for Low-Latency Inference at the Edge
by: Kafetzis, Dimitrios, et al.
Published: (2025)
by: Kafetzis, Dimitrios, et al.
Published: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
by: Chen, Xing, et al.
Published: (2025)
by: Chen, Xing, et al.
Published: (2025)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024)
by: Wang, Siqi, et al.
Published: (2024)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
by: Sheng, Jinhao, et al.
Published: (2025)
by: Sheng, Jinhao, et al.
Published: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
by: Wu, Tian, et al.
Published: (2025)
by: Wu, Tian, et al.
Published: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2023)
by: Jiang, Youhe, et al.
Published: (2023)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets
by: Liu, Guangyuan, et al.
Published: (2025)
by: Liu, Guangyuan, et al.
Published: (2025)
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Proactive and Reactive Autoscaling Techniques for Edge Computing
by: Gupta, Suhrid, et al.
Published: (2025)
by: Gupta, Suhrid, et al.
Published: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
by: Zhao, Yihao, et al.
Published: (2025)
by: Zhao, Yihao, et al.
Published: (2025)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
by: Zhang, Runhua, et al.
Published: (2025)
by: Zhang, Runhua, et al.
Published: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
by: Zhan, Huiyou, et al.
Published: (2025)
by: Zhan, Huiyou, et al.
Published: (2025)
Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
by: Moghadampanah, Mona, et al.
Published: (2025)
by: Moghadampanah, Mona, et al.
Published: (2025)
Federated Learning as a Service for Hierarchical Edge Networks with Heterogeneous Models
by: Gao, Wentao, et al.
Published: (2024)
by: Gao, Wentao, et al.
Published: (2024)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
by: Wolfrath, Joel, et al.
Published: (2025)
by: Wolfrath, Joel, et al.
Published: (2025)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
by: Wu, Feiyang, et al.
Published: (2025)
by: Wu, Feiyang, et al.
Published: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
by: Wang, Liujianfu, et al.
Published: (2025)
by: Wang, Liujianfu, et al.
Published: (2025)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Similar Items
-
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
by: Qiu, Houming, et al.
Published: (2025) -
Lightweight Federated Learning over Wireless Edge Networks
by: Hou, Xiangwang, et al.
Published: (2025) -
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
by: Zhu, Juan, et al.
Published: (2026) -
Approximated Coded Computing: Towards Fast, Private and Secure Distributed Machine Learning
by: Qiu, Houming, et al.
Published: (2024) -
Federated Unlearning in Edge Networks: A Survey of Fundamentals, Challenges, Practical Applications and Future Directions
by: Ng, Jer Shyuan, et al.
Published: (2026)