Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cheng, Jiaming, Nguyen, Duong Tung
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915924126203904
author Cheng, Jiaming
Nguyen, Duong Tung
author_facet Cheng, Jiaming
Nguyen, Duong Tung
contents This paper investigates the optimal allocation of large language model (LLM) inference workloads across heterogeneous edge data centers over time. Each data center features on-site renewable generation and faces dynamic electricity prices and spatiotemporal variability in renewable availability. We propose Green-LLM, a lexicographic multi-objective optimization framework that addresses this challenge without requiring manual weight tuning. The proposed model incorporates real-world constraints, including token-dependent processing delay and energy consumption, heterogeneous hardware capabilities, dynamic renewable generation, and spatiotemporal variations in electricity prices and carbon intensity. Unlike existing approaches that optimize individual environmental metrics in isolation, Green-LLM jointly minimizes operational cost, carbon emissions, and delay penalty while enforcing water consumption constraints to ensure both sustainability and quality-of-service requirements. Numerical results demonstrate that Green-LLM achieves significant reductions in carbon emissions and water consumption while maintaining operational costs within 3% of the minimum and ensuring sub-2-second response latency. These findings show that sustainable LLM inference can be achieved without sacrificing service quality or economic efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference
Cheng, Jiaming
Nguyen, Duong Tung
Networking and Internet Architecture
Distributed, Parallel, and Cluster Computing
Systems and Control
Optimization and Control
This paper investigates the optimal allocation of large language model (LLM) inference workloads across heterogeneous edge data centers over time. Each data center features on-site renewable generation and faces dynamic electricity prices and spatiotemporal variability in renewable availability. We propose Green-LLM, a lexicographic multi-objective optimization framework that addresses this challenge without requiring manual weight tuning. The proposed model incorporates real-world constraints, including token-dependent processing delay and energy consumption, heterogeneous hardware capabilities, dynamic renewable generation, and spatiotemporal variations in electricity prices and carbon intensity. Unlike existing approaches that optimize individual environmental metrics in isolation, Green-LLM jointly minimizes operational cost, carbon emissions, and delay penalty while enforcing water consumption constraints to ensure both sustainability and quality-of-service requirements. Numerical results demonstrate that Green-LLM achieves significant reductions in carbon emissions and water consumption while maintaining operational costs within 3% of the minimum and ensuring sub-2-second response latency. These findings show that sustainable LLM inference can be achieved without sacrificing service quality or economic efficiency.
title Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference
topic Networking and Internet Architecture
Distributed, Parallel, and Cluster Computing
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2507.09942