Benchmarking Energy Efficiency of Large Language Models Using vLLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pronk, K., Zhao, Q.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911149083066368
author Pronk, K.
Zhao, Q.
author_facet Pronk, K.
Zhao, Q.
contents The prevalence of Large Language Models (LLMs) is having an growing impact on the climate due to the substantial energy required for their deployment and use. To create awareness for developers who are implementing LLMs in their products, there is a strong need to collect more information about the energy efficiency of LLMs. While existing research has evaluated the energy efficiency of various models, these benchmarks often fall short of representing realistic production scenarios. In this paper, we introduce the LLM Efficiency Benchmark, designed to simulate real-world usage conditions. Our benchmark utilizes vLLM, a high-throughput, production-ready LLM serving backend that optimizes model performance and efficiency. We examine how factors such as model size, architecture, and concurrent request volume affect inference energy efficiency. Our findings demonstrate that it is possible to create energy efficiency benchmarks that better reflect practical deployment conditions, providing valuable insights for developers aiming to build more sustainable AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Energy Efficiency of Large Language Models Using vLLM
Pronk, K.
Zhao, Q.
Software Engineering
Artificial Intelligence
68T01
I.2.7
The prevalence of Large Language Models (LLMs) is having an growing impact on the climate due to the substantial energy required for their deployment and use. To create awareness for developers who are implementing LLMs in their products, there is a strong need to collect more information about the energy efficiency of LLMs. While existing research has evaluated the energy efficiency of various models, these benchmarks often fall short of representing realistic production scenarios. In this paper, we introduce the LLM Efficiency Benchmark, designed to simulate real-world usage conditions. Our benchmark utilizes vLLM, a high-throughput, production-ready LLM serving backend that optimizes model performance and efficiency. We examine how factors such as model size, architecture, and concurrent request volume affect inference energy efficiency. Our findings demonstrate that it is possible to create energy efficiency benchmarks that better reflect practical deployment conditions, providing valuable insights for developers aiming to build more sustainable AI systems.
title Benchmarking Energy Efficiency of Large Language Models Using vLLM
topic Software Engineering
Artificial Intelligence
68T01
I.2.7
url https://arxiv.org/abs/2509.08867