LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Faiz, Ahmad, Kaneda, Sotaro, Wang, Ruhan, Osi, Rita, Sharma, Prateek, Chen, Fan, Jiang, Lei
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916097399193600
author Faiz, Ahmad
Kaneda, Sotaro
Wang, Ruhan
Osi, Rita
Sharma, Prateek
Chen, Fan
Jiang, Lei
author_facet Faiz, Ahmad
Kaneda, Sotaro
Wang, Ruhan
Osi, Rita
Sharma, Prateek
Chen, Fan
Jiang, Lei
contents The carbon footprint associated with large language models (LLMs) is a significant concern, encompassing emissions from their training, inference, experimentation, and storage processes, including operational and embodied carbon emissions. An essential aspect is accurately estimating the carbon impact of emerging LLMs even before their training, which heavily relies on GPU usage. Existing studies have reported the carbon footprint of LLM training, but only one tool, mlco2, can predict the carbon footprint of new neural networks prior to physical training. However, mlco2 has several serious limitations. It cannot extend its estimation to dense or mixture-of-experts (MoE) LLMs, disregards critical architectural parameters, focuses solely on GPUs, and cannot model embodied carbon footprints. Addressing these gaps, we introduce \textit{\carb}, an end-to-end carbon footprint projection model designed for both dense and MoE LLMs. Compared to mlco2, \carb~significantly enhances the accuracy of carbon footprint estimations for various LLMs. The source code is released at \url{https://github.com/SotaroKaneda/MLCarbon}.
format Preprint
id arxiv_https___arxiv_org_abs_2309_14393
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
Faiz, Ahmad
Kaneda, Sotaro
Wang, Ruhan
Osi, Rita
Sharma, Prateek
Chen, Fan
Jiang, Lei
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
The carbon footprint associated with large language models (LLMs) is a significant concern, encompassing emissions from their training, inference, experimentation, and storage processes, including operational and embodied carbon emissions. An essential aspect is accurately estimating the carbon impact of emerging LLMs even before their training, which heavily relies on GPU usage. Existing studies have reported the carbon footprint of LLM training, but only one tool, mlco2, can predict the carbon footprint of new neural networks prior to physical training. However, mlco2 has several serious limitations. It cannot extend its estimation to dense or mixture-of-experts (MoE) LLMs, disregards critical architectural parameters, focuses solely on GPUs, and cannot model embodied carbon footprints. Addressing these gaps, we introduce \textit{\carb}, an end-to-end carbon footprint projection model designed for both dense and MoE LLMs. Compared to mlco2, \carb~significantly enhances the accuracy of carbon footprint estimations for various LLMs. The source code is released at \url{https://github.com/SotaroKaneda/MLCarbon}.
title LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2309.14393