Temperature-Aware Scheduling of LLM Inference in Large-Scale Geo-Distributed Edge Data Centers with Distributed Optimization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Khalatbarisoltani, Arash, Mahmoudi, Amin, Han, Jie, Saeed, Muhammad, Liu, Wenxue, Li, Jinwen, Kahourzade, Solmaz, Yazdani, Amirmehdi, Hu, Xiaosong |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Leveraging Quantum Annealing for Large-Scale Household Energy Scheduling with Hydrogen Storage
par: Khalatbarisoltani, Arash, et autres
Publié: (2026)
par: Khalatbarisoltani, Arash, et autres
Publié: (2026)
Driving Condition-Aware Multi-Agent Integrated Power and Thermal Management for Hybrid Electric Vehicles
par: Cui, Hanghang, et autres
Publié: (2026)
par: Cui, Hanghang, et autres
Publié: (2026)
Comparative Analysis of Control Observer-Based Methods for State Estimation of Lithium-Ion Batteries in Practical Scenarios
par: Saeed, Muhammad, et autres
Publié: (2024)
par: Saeed, Muhammad, et autres
Publié: (2024)
Traffic-aware Hierarchical Integrated Thermal and Energy Management for Connected HEVs
par: Han, Jie, et autres
Publié: (2026)
par: Han, Jie, et autres
Publié: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
par: Yuan, Renzhong, et autres
Publié: (2026)
par: Yuan, Renzhong, et autres
Publié: (2026)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
par: Zheng, Peirong, et autres
Publié: (2026)
par: Zheng, Peirong, et autres
Publié: (2026)
Argus: Token Aware Distributed LLM Inference Optimization
par: Wu, Panlong, et autres
Publié: (2025)
par: Wu, Panlong, et autres
Publié: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
par: Chen, Jiu, et autres
Publié: (2026)
par: Chen, Jiu, et autres
Publié: (2026)
PRISM: Distributed Inference for Foundation Models at Edge
par: Qazi, Muhammad Azlan, et autres
Publié: (2025)
par: Qazi, Muhammad Azlan, et autres
Publié: (2025)
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
par: Moore, Hayden, et autres
Publié: (2025)
par: Moore, Hayden, et autres
Publié: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
par: Zhang, Han, et autres
Publié: (2026)
par: Zhang, Han, et autres
Publié: (2026)
Priority-Aware Model-Distributed Inference at Edge Networks
par: Li, Teng, et autres
Publié: (2024)
par: Li, Teng, et autres
Publié: (2024)
Electricity Price-Aware Scheduling of Data Center Cooling
par: Khojaste, Arash, et autres
Publié: (2025)
par: Khojaste, Arash, et autres
Publié: (2025)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
par: Nguyen, Chanh, et autres
Publié: (2026)
par: Nguyen, Chanh, et autres
Publié: (2026)
Potential Distribution Theory of Alchemical Transfer
par: Azimi, Solmaz, et autres
Publié: (2024)
par: Azimi, Solmaz, et autres
Publié: (2024)
Submodular Maximization Subject to Uniform and Partition Matroids: From Theory to Practical Applications and Distributed Solutions
par: Kia, Solmaz S.
Publié: (2025)
par: Kia, Solmaz S.
Publié: (2025)
Dependency-Aware Online Caching
par: Dallot, Julien, et autres
Publié: (2024)
par: Dallot, Julien, et autres
Publié: (2024)
Differentially Private Distributed Inference
par: Papachristou, Marios, et autres
Publié: (2024)
par: Papachristou, Marios, et autres
Publié: (2024)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
par: Qazi, Muhammad Azlan, et autres
Publié: (2026)
par: Qazi, Muhammad Azlan, et autres
Publié: (2026)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
par: Kasaei, Seyed Amir, et autres
Publié: (2025)
par: Kasaei, Seyed Amir, et autres
Publié: (2025)
DILEMMA: Joint LLM Quantization and Distributed LLM Inference Over Edge Computing Systems
par: Hosseinzadeh, Minoo, et autres
Publié: (2025)
par: Hosseinzadeh, Minoo, et autres
Publié: (2025)
Privacy-Aware Multi-Device Cooperative Edge Inference with Distributed Resource Bidding
par: Zhuang, Wenhao, et autres
Publié: (2024)
par: Zhuang, Wenhao, et autres
Publié: (2024)
Channel Capacity-Aware Distributed Encoding for Multi-View Sensing and Edge Inference
par: Yang, Mingjie, et autres
Publié: (2024)
par: Yang, Mingjie, et autres
Publié: (2024)
AgentFM: Role-Aware Failure Management for Distributed Databases with LLM-Driven Multi-Agents
par: Zhang, Lingzhe, et autres
Publié: (2025)
par: Zhang, Lingzhe, et autres
Publié: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Multi-agent Coverage Control: From Discrete Assignments to Continuous Multi-agent Distribution Matching
par: Kia, Solmaz, et autres
Publié: (2024)
par: Kia, Solmaz, et autres
Publié: (2024)
Dynamic Energy-Aware Task Scheduling using Real-Time Resource Monitoring in Distributed Edge Environments
par: Vanapalli Jayanth Sai, et autres
Publié: (2026)
par: Vanapalli Jayanth Sai, et autres
Publié: (2026)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
par: Yang, Zheming, et autres
Publié: (2024)
par: Yang, Zheming, et autres
Publié: (2024)
Communication-Efficient Multi-Modal Edge Inference via Uncertainty-Aware Distributed Learning
par: Zhao, Hang, et autres
Publié: (2026)
par: Zhao, Hang, et autres
Publié: (2026)
Optimizing a Straight‐Bladed Vertical Axis Wind Turbine With Computational Fluid Dynamics (CFD), Artificial Neural Network (ANN), and Genetic Algorithm (GA)
par: Sepehr Sanaye, et autres
Publié: (2025)
par: Sepehr Sanaye, et autres
Publié: (2025)
Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference
par: Cheng, Jiaming, et autres
Publié: (2025)
par: Cheng, Jiaming, et autres
Publié: (2025)
SLA-Aware Distributed LLM Inference Across Device-RAN-Cloud
par: Yet, Hariz, et autres
Publié: (2026)
par: Yet, Hariz, et autres
Publié: (2026)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
par: Deng, Xiumei, et autres
Publié: (2025)
par: Deng, Xiumei, et autres
Publié: (2025)
DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
par: Xu, Yaodan, et autres
Publié: (2026)
par: Xu, Yaodan, et autres
Publié: (2026)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
par: Wolfrath, Joel, et autres
Publié: (2025)
par: Wolfrath, Joel, et autres
Publié: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
par: Karfakis, George, et autres
Publié: (2025)
par: Karfakis, George, et autres
Publié: (2025)
CAFE: Carbon-Aware Federated Learning in Geographically Distributed Data Centers
par: Bian, Jieming, et autres
Publié: (2023)
par: Bian, Jieming, et autres
Publié: (2023)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
par: Oh, Hyungjun, et autres
Publié: (2024)
par: Oh, Hyungjun, et autres
Publié: (2024)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
par: Zhang, Ziyang, et autres
Publié: (2025)
par: Zhang, Ziyang, et autres
Publié: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
par: Rajashekar, Kolichala, et autres
Publié: (2025)
par: Rajashekar, Kolichala, et autres
Publié: (2025)
Documents similaires
-
Leveraging Quantum Annealing for Large-Scale Household Energy Scheduling with Hydrogen Storage
par: Khalatbarisoltani, Arash, et autres
Publié: (2026) -
Driving Condition-Aware Multi-Agent Integrated Power and Thermal Management for Hybrid Electric Vehicles
par: Cui, Hanghang, et autres
Publié: (2026) -
Comparative Analysis of Control Observer-Based Methods for State Estimation of Lithium-Ion Batteries in Practical Scenarios
par: Saeed, Muhammad, et autres
Publié: (2024) -
Traffic-aware Hierarchical Integrated Thermal and Energy Management for Connected HEVs
par: Han, Jie, et autres
Publié: (2026) -
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
par: Yuan, Renzhong, et autres
Publié: (2026)