Saved in:
Bibliographic Details
Main Author: Almog, Tom
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.13516
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912590117994496
author Almog, Tom
author_facet Almog, Tom
contents As machine learning models grow increasingly complex and computationally demanding, understanding the environmental impact of training decisions becomes critical for sustainable AI development. This paper presents a comprehensive empirical study investigating the relationship between optimizer choice and energy efficiency in neural network training. We conducted 360 controlled experiments across three benchmark datasets (MNIST, CIFAR-10, CIFAR-100) using eight popular optimizers (SGD, Adam, AdamW, RMSprop, Adagrad, Adadelta, Adamax, NAdam) with 15 random seeds each. Using CodeCarbon for precise energy tracking on Apple M1 Pro hardware, we measured training duration, peak memory usage, carbon dioxide emissions, and final model performance. Our findings reveal substantial trade-offs between training speed, accuracy, and environmental impact that vary across datasets and model complexity. We identify AdamW and NAdam as consistently efficient choices, while SGD demonstrates superior performance on complex datasets despite higher emissions. These results provide actionable insights for practitioners seeking to balance performance and sustainability in machine learning workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13516
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Analysis of Optimizer Choice on Energy Efficiency and Performance in Neural Network Training
Almog, Tom
Machine Learning
68T05 (Primary) 90C30, 68W40 (Secondary)
As machine learning models grow increasingly complex and computationally demanding, understanding the environmental impact of training decisions becomes critical for sustainable AI development. This paper presents a comprehensive empirical study investigating the relationship between optimizer choice and energy efficiency in neural network training. We conducted 360 controlled experiments across three benchmark datasets (MNIST, CIFAR-10, CIFAR-100) using eight popular optimizers (SGD, Adam, AdamW, RMSprop, Adagrad, Adadelta, Adamax, NAdam) with 15 random seeds each. Using CodeCarbon for precise energy tracking on Apple M1 Pro hardware, we measured training duration, peak memory usage, carbon dioxide emissions, and final model performance. Our findings reveal substantial trade-offs between training speed, accuracy, and environmental impact that vary across datasets and model complexity. We identify AdamW and NAdam as consistently efficient choices, while SGD demonstrates superior performance on complex datasets despite higher emissions. These results provide actionable insights for practitioners seeking to balance performance and sustainability in machine learning workflows.
title An Analysis of Optimizer Choice on Energy Efficiency and Performance in Neural Network Training
topic Machine Learning
68T05 (Primary) 90C30, 68W40 (Secondary)
url https://arxiv.org/abs/2509.13516