Carbon-Aware Quality Adaptation for Energy-Intensive Services

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wiesner, Philipp, Grinwald, Dennis, Weiß, Philipp, Wilhelm, Patrick, Khalili, Ramin, Kao, Odej
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912942287486976
author Wiesner, Philipp
Grinwald, Dennis
Weiß, Philipp
Wilhelm, Patrick
Khalili, Ramin
Kao, Odej
author_facet Wiesner, Philipp
Grinwald, Dennis
Weiß, Philipp
Wilhelm, Patrick
Khalili, Ramin
Kao, Odej
contents The energy demand of modern cloud services, particularly those related to generative AI, is increasing at an unprecedented pace. To date, carbon-aware computing strategies have primarily focused on batch process scheduling or geo-distributed load balancing. However, such approaches are not applicable to services that require constant availability at specific locations due to latency, privacy, data, or infrastructure constraints. In this paper, we explore how the carbon footprint of energy-intensive services can be reduced by adjusting the fraction of requests served by different service quality tiers. We show that adapting this quality of responses with respect to grid carbon intensity can lead to additional carbon savings beyond resource and energy efficiency. Building on this, we introduce a forecast-based multi-horizon optimization that reaches close-to-optimal carbon savings and is able to automatically adapt service quality for best-effort users to stay within an annual carbon budget. Our approach can reduce the emissions of large-scale LLM services, which we estimate at multiple 10,000 tons of CO2 annually, by up to 10%.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19058
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Carbon-Aware Quality Adaptation for Energy-Intensive Services
Wiesner, Philipp
Grinwald, Dennis
Weiß, Philipp
Wilhelm, Patrick
Khalili, Ramin
Kao, Odej
Distributed, Parallel, and Cluster Computing
Systems and Control
The energy demand of modern cloud services, particularly those related to generative AI, is increasing at an unprecedented pace. To date, carbon-aware computing strategies have primarily focused on batch process scheduling or geo-distributed load balancing. However, such approaches are not applicable to services that require constant availability at specific locations due to latency, privacy, data, or infrastructure constraints. In this paper, we explore how the carbon footprint of energy-intensive services can be reduced by adjusting the fraction of requests served by different service quality tiers. We show that adapting this quality of responses with respect to grid carbon intensity can lead to additional carbon savings beyond resource and energy efficiency. Building on this, we introduce a forecast-based multi-horizon optimization that reaches close-to-optimal carbon savings and is able to automatically adapt service quality for best-effort users to stay within an annual carbon budget. Our approach can reduce the emissions of large-scale LLM services, which we estimate at multiple 10,000 tons of CO2 annually, by up to 10%.
title Carbon-Aware Quality Adaptation for Energy-Intensive Services
topic Distributed, Parallel, and Cluster Computing
Systems and Control
url https://arxiv.org/abs/2411.19058