Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dutt, Anurag, Choi, Young Won, Sil, Avirup, Gandhi, Anshul, Balasubramanian, Aruna, Balasubramanian, Niranjan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912764301148160
author Dutt, Anurag
Choi, Young Won
Sil, Avirup
Gandhi, Anshul
Balasubramanian, Aruna
Balasubramanian, Niranjan
author_facet Dutt, Anurag
Choi, Young Won
Sil, Avirup
Gandhi, Anshul
Balasubramanian, Aruna
Balasubramanian, Niranjan
contents With the widespread adoption of Large Language Models (LLMs), energy costs of running LLMs is quickly becoming a critical concern. However, precisely measuring the energy consumption of LLMs is often infeasible because hardware-based power monitors are not always accessible and software-based energy measurement tools are not accurate. While various prediction techniques have been developed to estimate LLM energy consumption, these approaches are limited to single-GPU environments and thus are not applicable to modern LLM inference which is typically parallelized across multiple GPUs. In this work, we remedy this gap and introduce PIE-P, a fine-grained energy prediction framework for multi-GPU inference, including tensor, pipeline, and data parallelism. Predicting the energy under parallelized inference is complicated by the non-determinism in inter-GPU communication, additional communication overheads, and difficulties in isolating energy during the communication/synchronization phase. We develop a scalable prediction framework that addresses these issues via precise sampling, fine-grained modeling of inter-GPU communication, and careful accounting of parallelization overhead. Our evaluation results show that PIE-P yields accurate and fine-grained energy predictions across parallelism strategies, significantly outperforming baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12801
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
Dutt, Anurag
Choi, Young Won
Sil, Avirup
Gandhi, Anshul
Balasubramanian, Aruna
Balasubramanian, Niranjan
Distributed, Parallel, and Cluster Computing
Performance
With the widespread adoption of Large Language Models (LLMs), energy costs of running LLMs is quickly becoming a critical concern. However, precisely measuring the energy consumption of LLMs is often infeasible because hardware-based power monitors are not always accessible and software-based energy measurement tools are not accurate. While various prediction techniques have been developed to estimate LLM energy consumption, these approaches are limited to single-GPU environments and thus are not applicable to modern LLM inference which is typically parallelized across multiple GPUs. In this work, we remedy this gap and introduce PIE-P, a fine-grained energy prediction framework for multi-GPU inference, including tensor, pipeline, and data parallelism. Predicting the energy under parallelized inference is complicated by the non-determinism in inter-GPU communication, additional communication overheads, and difficulties in isolating energy during the communication/synchronization phase. We develop a scalable prediction framework that addresses these issues via precise sampling, fine-grained modeling of inter-GPU communication, and careful accounting of parallelization overhead. Our evaluation results show that PIE-P yields accurate and fine-grained energy predictions across parallelism strategies, significantly outperforming baselines.
title Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2512.12801