Towards Green AI: Energy-Efficient Training and Inference of Large Language Models
Fuente:
Zenodo
Gespeichert in:
| Hauptverfasser: | , |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2025
|
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901550654291968 |
|---|---|
| author | Vats, Harshwardhan Kumar, Ritesh |
| author_facet | Vats, Harshwardhan Kumar, Ritesh |
| contents | <p>Abstract:<br>Large Language Models (LLMs) have transformed natural language processing with their<br>remarkable capabilities, yet their ever-growing computational demands come with steep<br>energy costs and increasing carbon emissions. This paper provides an in-depth,<br>comparative case study titled “Towards Green AI: Energy-Efficient Training and<br>Inference of LLMs,” examining how efficiency-oriented methods can make large-scale<br>AI more sustainable. Using a qualitative, analytical approach, it draws on research papers,<br>benchmarks, and technical reports published between 2020 and 2025 to highlight key<br>strategies for improving energy efficiency—such as low-precision computation<br>(FP16/FP8), parameter-efficient fine-tuning (LoRA), low-bit quantized inference<br>(ATOM), optimized attention kernels (FlashAttention-2), and inference accelerations like<br>speculative decoding. The study discusses the practical trade-offs of these techniques,<br>outlines best practices for measuring efficiency, and explores their broader implications<br>for developing and deploying greener AI systems at scale.</p> <p><br>Keywords: Green AI, case study analysis, energy efficiency, large language models,<br>quantization, FP8, LoRA</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17652598 |
| institution | Zenodo |
| language | eng |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Towards Green AI: Energy-Efficient Training and Inference of Large Language Models Vats, Harshwardhan Kumar, Ritesh <p>Abstract:<br>Large Language Models (LLMs) have transformed natural language processing with their<br>remarkable capabilities, yet their ever-growing computational demands come with steep<br>energy costs and increasing carbon emissions. This paper provides an in-depth,<br>comparative case study titled “Towards Green AI: Energy-Efficient Training and<br>Inference of LLMs,” examining how efficiency-oriented methods can make large-scale<br>AI more sustainable. Using a qualitative, analytical approach, it draws on research papers,<br>benchmarks, and technical reports published between 2020 and 2025 to highlight key<br>strategies for improving energy efficiency—such as low-precision computation<br>(FP16/FP8), parameter-efficient fine-tuning (LoRA), low-bit quantized inference<br>(ATOM), optimized attention kernels (FlashAttention-2), and inference accelerations like<br>speculative decoding. The study discusses the practical trade-offs of these techniques,<br>outlines best practices for measuring efficiency, and explores their broader implications<br>for developing and deploying greener AI systems at scale.</p> <p><br>Keywords: Green AI, case study analysis, energy efficiency, large language models,<br>quantization, FP8, LoRA</p> |
| title | Towards Green AI: Energy-Efficient Training and Inference of Large Language Models |
| url | https://doi.org/10.5281/zenodo.17652598 |