Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wilkins, Grant, Keshav, Srinivasan, Mortier, Richard
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913409094647808
author Wilkins, Grant
Keshav, Srinivasan
Mortier, Richard
author_facet Wilkins, Grant
Keshav, Srinivasan
Mortier, Richard
contents Both the training and use of Large Language Models (LLMs) require large amounts of energy. Their increasing popularity, therefore, raises critical concerns regarding the energy efficiency and sustainability of data centers that host them. This paper addresses the challenge of reducing energy consumption in data centers running LLMs. We propose a hybrid data center model that uses a cost-based scheduling framework to dynamically allocate LLM tasks across hardware accelerators that differ in their energy efficiencies and computational capabilities. Specifically, our workload-aware strategy determines whether tasks are processed on energy-efficient processors or high-performance GPUs based on the number of input and output tokens in a query. Our analysis of a representative LLM dataset, finds that this hybrid strategy can reduce CPU+GPU energy consumption by 7.5% compared to a workload-unaware baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00010
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
Wilkins, Grant
Keshav, Srinivasan
Mortier, Richard
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Both the training and use of Large Language Models (LLMs) require large amounts of energy. Their increasing popularity, therefore, raises critical concerns regarding the energy efficiency and sustainability of data centers that host them. This paper addresses the challenge of reducing energy consumption in data centers running LLMs. We propose a hybrid data center model that uses a cost-based scheduling framework to dynamically allocate LLM tasks across hardware accelerators that differ in their energy efficiencies and computational capabilities. Specifically, our workload-aware strategy determines whether tasks are processed on energy-efficient processors or high-performance GPUs based on the number of input and output tokens in a query. Our analysis of a representative LLM dataset, finds that this hybrid strategy can reduce CPU+GPU energy consumption by 7.5% compared to a workload-unaware baseline.
title Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2407.00010