Benchmarking Neural Network Training Algorithms

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dahl, George E., Schneider, Frank, Nado, Zachary, Agarwal, Naman, Sastry, Chandramouli Shama, Hennig, Philipp, Medapati, Sourabh, Eschenhagen, Runa, Kasimbeg, Priya, Suo, Daniel, Bae, Juhan, Gilmer, Justin, Peirson, Abel L., Khan, Bilal, Anil, Rohan, Rabbat, Mike, Krishnan, Shankar, Snider, Daniel, Amid, Ehsan, Chen, Kongtao, Maddison, Chris J., Vasudev, Rakshith, Badura, Michal, Garg, Ankush, Mattson, Peter
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916798065016832
author Dahl, George E.
Schneider, Frank
Nado, Zachary
Agarwal, Naman
Sastry, Chandramouli Shama
Hennig, Philipp
Medapati, Sourabh
Eschenhagen, Runa
Kasimbeg, Priya
Suo, Daniel
Bae, Juhan
Gilmer, Justin
Peirson, Abel L.
Khan, Bilal
Anil, Rohan
Rabbat, Mike
Krishnan, Shankar
Snider, Daniel
Amid, Ehsan
Chen, Kongtao
Maddison, Chris J.
Vasudev, Rakshith
Badura, Michal
Garg, Ankush
Mattson, Peter
author_facet Dahl, George E.
Schneider, Frank
Nado, Zachary
Agarwal, Naman
Sastry, Chandramouli Shama
Hennig, Philipp
Medapati, Sourabh
Eschenhagen, Runa
Kasimbeg, Priya
Suo, Daniel
Bae, Juhan
Gilmer, Justin
Peirson, Abel L.
Khan, Bilal
Anil, Rohan
Rabbat, Mike
Krishnan, Shankar
Snider, Daniel
Amid, Ehsan
Chen, Kongtao
Maddison, Chris J.
Vasudev, Rakshith
Badura, Michal
Garg, Ankush
Mattson, Peter
contents Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workloads (e.g., better update rules, tuning protocols, learning rate schedules, or data selection schemes) could save time, save computational resources, and lead to better, more accurate, models. Unfortunately, as a community, we are currently unable to reliably identify training algorithm improvements, or even determine the state-of-the-art training algorithm. In this work, using concrete experiments, we argue that real progress in speeding up training requires new benchmarks that resolve three basic challenges faced by empirical comparisons of training algorithms: (1) how to decide when training is complete and precisely measure training time, (2) how to handle the sensitivity of measurements to exact workload details, and (3) how to fairly compare algorithms that require hyperparameter tuning. In order to address these challenges, we introduce a new, competitive, time-to-result benchmark using multiple workloads running on fixed hardware, the AlgoPerf: Training Algorithms benchmark. Our benchmark includes a set of workload variants that make it possible to detect benchmark submissions that are more robust to workload changes than current widely-used methods. Finally, we evaluate baseline submissions constructed using various optimizers that represent current practice, as well as other optimizers that have recently received attention in the literature. These baseline results collectively demonstrate the feasibility of our benchmark, show that non-trivial gaps between methods exist, and set a provisional state-of-the-art for future benchmark submissions to try and surpass.
format Preprint
id arxiv_https___arxiv_org_abs_2306_07179
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Benchmarking Neural Network Training Algorithms
Dahl, George E.
Schneider, Frank
Nado, Zachary
Agarwal, Naman
Sastry, Chandramouli Shama
Hennig, Philipp
Medapati, Sourabh
Eschenhagen, Runa
Kasimbeg, Priya
Suo, Daniel
Bae, Juhan
Gilmer, Justin
Peirson, Abel L.
Khan, Bilal
Anil, Rohan
Rabbat, Mike
Krishnan, Shankar
Snider, Daniel
Amid, Ehsan
Chen, Kongtao
Maddison, Chris J.
Vasudev, Rakshith
Badura, Michal
Garg, Ankush
Mattson, Peter
Machine Learning
Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workloads (e.g., better update rules, tuning protocols, learning rate schedules, or data selection schemes) could save time, save computational resources, and lead to better, more accurate, models. Unfortunately, as a community, we are currently unable to reliably identify training algorithm improvements, or even determine the state-of-the-art training algorithm. In this work, using concrete experiments, we argue that real progress in speeding up training requires new benchmarks that resolve three basic challenges faced by empirical comparisons of training algorithms: (1) how to decide when training is complete and precisely measure training time, (2) how to handle the sensitivity of measurements to exact workload details, and (3) how to fairly compare algorithms that require hyperparameter tuning. In order to address these challenges, we introduce a new, competitive, time-to-result benchmark using multiple workloads running on fixed hardware, the AlgoPerf: Training Algorithms benchmark. Our benchmark includes a set of workload variants that make it possible to detect benchmark submissions that are more robust to workload changes than current widely-used methods. Finally, we evaluate baseline submissions constructed using various optimizers that represent current practice, as well as other optimizers that have recently received attention in the literature. These baseline results collectively demonstrate the feasibility of our benchmark, show that non-trivial gaps between methods exist, and set a provisional state-of-the-art for future benchmark submissions to try and surpass.
title Benchmarking Neural Network Training Algorithms
topic Machine Learning
url https://arxiv.org/abs/2306.07179