Towards Green AI: Energy-Efficient Training and Inference of Large Language Models

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vats, Harshwardhan, Kumar, Ritesh
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2025
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901550654291968
author Vats, Harshwardhan
Kumar, Ritesh
author_facet Vats, Harshwardhan
Kumar, Ritesh
contents <p>Abstract:<br>Large Language Models (LLMs) have transformed natural language processing with their<br>remarkable capabilities, yet their ever-growing computational demands come with steep<br>energy costs and increasing carbon emissions. This paper provides an in-depth,<br>comparative case study titled “Towards Green AI: Energy-Efficient Training and<br>Inference of LLMs,” examining how efficiency-oriented methods can make large-scale<br>AI more sustainable. Using a qualitative, analytical approach, it draws on research papers,<br>benchmarks, and technical reports published between 2020 and 2025 to highlight key<br>strategies for improving energy efficiency—such as low-precision computation<br>(FP16/FP8), parameter-efficient fine-tuning (LoRA), low-bit quantized inference<br>(ATOM), optimized attention kernels (FlashAttention-2), and inference accelerations like<br>speculative decoding. The study discusses the practical trade-offs of these techniques,<br>outlines best practices for measuring efficiency, and explores their broader implications<br>for developing and deploying greener AI systems at scale.</p> <p><br>Keywords: Green AI, case study analysis, energy efficiency, large language models,<br>quantization, FP8, LoRA</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17652598
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Towards Green AI: Energy-Efficient Training and Inference of Large Language Models
Vats, Harshwardhan
Kumar, Ritesh
<p>Abstract:<br>Large Language Models (LLMs) have transformed natural language processing with their<br>remarkable capabilities, yet their ever-growing computational demands come with steep<br>energy costs and increasing carbon emissions. This paper provides an in-depth,<br>comparative case study titled “Towards Green AI: Energy-Efficient Training and<br>Inference of LLMs,” examining how efficiency-oriented methods can make large-scale<br>AI more sustainable. Using a qualitative, analytical approach, it draws on research papers,<br>benchmarks, and technical reports published between 2020 and 2025 to highlight key<br>strategies for improving energy efficiency—such as low-precision computation<br>(FP16/FP8), parameter-efficient fine-tuning (LoRA), low-bit quantized inference<br>(ATOM), optimized attention kernels (FlashAttention-2), and inference accelerations like<br>speculative decoding. The study discusses the practical trade-offs of these techniques,<br>outlines best practices for measuring efficiency, and explores their broader implications<br>for developing and deploying greener AI systems at scale.</p> <p><br>Keywords: Green AI, case study analysis, energy efficiency, large language models,<br>quantization, FP8, LoRA</p>
title Towards Green AI: Energy-Efficient Training and Inference of Large Language Models
url https://doi.org/10.5281/zenodo.17652598