Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Özcan, Miray, Wiesner, Philipp, Weiß, Philipp, Kao, Odej
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909690098614272
author Özcan, Miray
Wiesner, Philipp
Weiß, Philipp
Kao, Odej
author_facet Özcan, Miray
Wiesner, Philipp
Weiß, Philipp
Kao, Odej
contents The environmental impact of Large Language Models (LLMs) is rising significantly, with inference now accounting for more than half of their total lifecycle carbon emissions. However, existing simulation frameworks, which are increasingly used to determine efficient LLM deployments, lack any concept of power and, therefore, cannot accurately estimate inference-related emissions. We present a simulation framework to assess the energy and carbon implications of LLM inference under varying deployment setups. First, we extend a high-fidelity LLM inference simulator with a GPU power model that estimates power consumption based on utilization metrics, enabling analysis across configurations like batch size, sequence length, and model parallelism. Second, we integrate simulation outputs into an energy system co-simulation environment to quantify carbon emissions under specific grid conditions and explore the potential of carbon-aware scheduling. Through scenario-based analysis, our framework reveals how inference parameters affect energy demand and carbon footprint, demonstrates a renewable offset potential of up to 69.2% in an illustrative deployment case, and provides a foundation for future carbon-aware inference infrastructure design.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
Özcan, Miray
Wiesner, Philipp
Weiß, Philipp
Kao, Odej
Distributed, Parallel, and Cluster Computing
The environmental impact of Large Language Models (LLMs) is rising significantly, with inference now accounting for more than half of their total lifecycle carbon emissions. However, existing simulation frameworks, which are increasingly used to determine efficient LLM deployments, lack any concept of power and, therefore, cannot accurately estimate inference-related emissions. We present a simulation framework to assess the energy and carbon implications of LLM inference under varying deployment setups. First, we extend a high-fidelity LLM inference simulator with a GPU power model that estimates power consumption based on utilization metrics, enabling analysis across configurations like batch size, sequence length, and model parallelism. Second, we integrate simulation outputs into an energy system co-simulation environment to quantify carbon emissions under specific grid conditions and explore the potential of carbon-aware scheduling. Through scenario-based analysis, our framework reveals how inference parameters affect energy demand and carbon footprint, demonstrates a renewable offset potential of up to 69.2% in an illustrative deployment case, and provides a foundation for future carbon-aware inference infrastructure design.
title Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2507.11417