Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
Fuente:
arXiv
Guardado en:
| Autores principales: | Moore, Hayden, Qi, Sirui, Hogade, Ninad, Milojicic, Dejan, Bash, Cullen, Pasricha, Sudeep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
por: Qi, Sirui, et al.
Publicado: (2024)
por: Qi, Sirui, et al.
Publicado: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
por: Qi, S., et al.
Publicado: (2024)
por: Qi, S., et al.
Publicado: (2024)
MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
por: Moore, H., et al.
Publicado: (2026)
por: Moore, H., et al.
Publicado: (2026)
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
por: Hogade, Ninad, et al.
Publicado: (2024)
por: Hogade, Ninad, et al.
Publicado: (2024)
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
por: Ramicetty, P., et al.
Publicado: (2026)
por: Ramicetty, P., et al.
Publicado: (2026)
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
por: Kamatar, Alok, et al.
Publicado: (2024)
por: Kamatar, Alok, et al.
Publicado: (2024)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
por: Rodrigues, Elvis, et al.
Publicado: (2025)
por: Rodrigues, Elvis, et al.
Publicado: (2025)
Hotspot-Aware Scheduling of Virtual Machines with Overcommitment for Ultimate Utilization in Cloud Datacenters
por: Wu, Jiaxi, et al.
Publicado: (2026)
por: Wu, Jiaxi, et al.
Publicado: (2026)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
por: Zhang, Han, et al.
Publicado: (2026)
por: Zhang, Han, et al.
Publicado: (2026)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
por: Chen, Tiancheng, et al.
Publicado: (2025)
por: Chen, Tiancheng, et al.
Publicado: (2025)
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
por: Jiang, Yankai, et al.
Publicado: (2024)
por: Jiang, Yankai, et al.
Publicado: (2024)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
por: Kumari, Suchi, et al.
Publicado: (2025)
por: Kumari, Suchi, et al.
Publicado: (2025)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
por: Hanafy, Walid A., et al.
Publicado: (2025)
por: Hanafy, Walid A., et al.
Publicado: (2025)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
por: Russo, Gabriele Russo, et al.
Publicado: (2024)
por: Russo, Gabriele Russo, et al.
Publicado: (2024)
Capsule: Efficient Player Isolation for Datacenters
por: Du, Zhouheng, et al.
Publicado: (2025)
por: Du, Zhouheng, et al.
Publicado: (2025)
Uncertainty-Aware Decarbonization for Datacenters
por: Li, Amy, et al.
Publicado: (2024)
por: Li, Amy, et al.
Publicado: (2024)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
por: Lettich, Francesco, et al.
Publicado: (2024)
por: Lettich, Francesco, et al.
Publicado: (2024)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2026)
por: Yang, Zheming, et al.
Publicado: (2026)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
por: Niewenhuis, Dante, et al.
Publicado: (2026)
por: Niewenhuis, Dante, et al.
Publicado: (2026)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
por: Nicolae, Radu, et al.
Publicado: (2026)
por: Nicolae, Radu, et al.
Publicado: (2026)
Tetris: Efficient Intra-Datacenter Calls Packing for Large Conferencing Services
por: Gandhi, Rohan, et al.
Publicado: (2025)
por: Gandhi, Rohan, et al.
Publicado: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
por: Peng, You, et al.
Publicado: (2026)
por: Peng, You, et al.
Publicado: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
por: Mehboob, Talha, et al.
Publicado: (2025)
por: Mehboob, Talha, et al.
Publicado: (2025)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
por: Yang, Dingyu, et al.
Publicado: (2026)
por: Yang, Dingyu, et al.
Publicado: (2026)
Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows
por: Schweisgut, Dominik, et al.
Publicado: (2026)
por: Schweisgut, Dominik, et al.
Publicado: (2026)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
por: Xue, Chunyu, et al.
Publicado: (2026)
por: Xue, Chunyu, et al.
Publicado: (2026)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
Distribution and Management of Datacenter Load Decoupling
por: Lin, Liuzixuan, et al.
Publicado: (2025)
por: Lin, Liuzixuan, et al.
Publicado: (2025)
Datacenter Energy Optimized Power Profiles
por: Narayanaswamy, Sreedhar, et al.
Publicado: (2025)
por: Narayanaswamy, Sreedhar, et al.
Publicado: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
por: Stojkovic, Jovan, et al.
Publicado: (2025)
por: Stojkovic, Jovan, et al.
Publicado: (2025)
Task Scheduling in Geo-Distributed Computing: A Survey
por: Wu, Yujian, et al.
Publicado: (2025)
por: Wu, Yujian, et al.
Publicado: (2025)
Carbon-Aware Workflow Scheduling with Fixed Mapping and Deadline Constraint
por: Schweisgut, Dominik, et al.
Publicado: (2025)
por: Schweisgut, Dominik, et al.
Publicado: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
por: Zhu, Botao, et al.
Publicado: (2025)
por: Zhu, Botao, et al.
Publicado: (2025)
FATE: Future-State-Aware Scheduling for Heterogeneous LLM Workflows
por: Huang, Zirui, et al.
Publicado: (2026)
por: Huang, Zirui, et al.
Publicado: (2026)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
por: Siavashi, Mohammad, et al.
Publicado: (2025)
por: Siavashi, Mohammad, et al.
Publicado: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
por: Wang, Qipeng
Publicado: (2026)
por: Wang, Qipeng
Publicado: (2026)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
por: Huang, Shaoyuan, et al.
Publicado: (2024)
por: Huang, Shaoyuan, et al.
Publicado: (2024)
Ejemplares similares
-
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
por: Qi, Sirui, et al.
Publicado: (2024) -
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
por: Qi, S., et al.
Publicado: (2024) -
MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
por: Moore, H., et al.
Publicado: (2026) -
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
por: Hogade, Ninad, et al.
Publicado: (2024) -
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
por: Ramicetty, P., et al.
Publicado: (2026)