Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Maliakel, Paul Joe, Ilager, Shashikant, Brandic, Ivona
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912921533022208
author Maliakel, Paul Joe
Ilager, Shashikant
Brandic, Ivona
author_facet Maliakel, Paul Joe
Ilager, Shashikant
Brandic, Ivona
contents LLM inference exhibits substantial variability across queries and execution phases, yet inference configurations are often applied uniformly. We present a measurement-driven characterization of workload heterogeneity and energy-performance behavior of LLM inference under GPU dynamic voltage and frequency scaling (DVFS). We evaluate five decoder-only LLMs (1B-32B parameters) across four NLP benchmarks using a controlled offline setup. We show that lightweight semantic features predict inference difficulty better than input length, with 44.5% of queries achieving comparable quality across model sizes. At the hardware level, the decode phase dominates inference time (77-91%) and is largely insensitive to GPU frequency. Consequently, reducing GPU frequency from 2842 MHz to 180 MHz achieves an average of 42% energy savings with only a 1-6% latency increase. We further provide a use case with an upper-bound analysis of the potential benefits of combining workload-aware model selection with phase-aware DVFS, motivating future energy-efficient LLM inference systems.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
Maliakel, Paul Joe
Ilager, Shashikant
Brandic, Ivona
Machine Learning
LLM inference exhibits substantial variability across queries and execution phases, yet inference configurations are often applied uniformly. We present a measurement-driven characterization of workload heterogeneity and energy-performance behavior of LLM inference under GPU dynamic voltage and frequency scaling (DVFS). We evaluate five decoder-only LLMs (1B-32B parameters) across four NLP benchmarks using a controlled offline setup. We show that lightweight semantic features predict inference difficulty better than input length, with 44.5% of queries achieving comparable quality across model sizes. At the hardware level, the decode phase dominates inference time (77-91%) and is largely insensitive to GPU frequency. Consequently, reducing GPU frequency from 2842 MHz to 180 MHz achieves an average of 42% energy savings with only a 1-6% latency increase. We further provide a use case with an upper-bound analysis of the potential benefits of combining workload-aware model selection with phase-aware DVFS, motivating future energy-efficient LLM inference systems.
title Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
topic Machine Learning
url https://arxiv.org/abs/2501.08219