How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Erben, Alexander, Mayer, Ruben, Jacobsen, Hans-Arno
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914818105016320
author Erben, Alexander
Mayer, Ruben
Jacobsen, Hans-Arno
author_facet Erben, Alexander
Mayer, Ruben
Jacobsen, Hans-Arno
contents This paper aims to answer the question: Can deep learning models be cost-efficiently trained on a global market of spot VMs spanning different data centers and cloud providers? To provide guidance, we extensively evaluate the cost and throughput implications of training in different zones, continents, and clouds for representative CV, NLP, and ASR models. To expand the current training options further, we compare the scalability potential for hybrid-cloud scenarios by adding cloud resources to on-premise hardware to improve training throughput. Finally, we show how leveraging spot instance pricing enables a new cost-efficient way to train models with multiple cheap VMs, trumping both more centralized and powerful hardware and even on-demand cloud offerings at competitive prices.
format Preprint
id arxiv_https___arxiv_org_abs_2306_03163
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study
Erben, Alexander
Mayer, Ruben
Jacobsen, Hans-Arno
Machine Learning
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
Performance
I.2.11; C.2.4; C.4; D.2.8
This paper aims to answer the question: Can deep learning models be cost-efficiently trained on a global market of spot VMs spanning different data centers and cloud providers? To provide guidance, we extensively evaluate the cost and throughput implications of training in different zones, continents, and clouds for representative CV, NLP, and ASR models. To expand the current training options further, we compare the scalability potential for hybrid-cloud scenarios by adding cloud resources to on-premise hardware to improve training throughput. Finally, we show how leveraging spot instance pricing enables a new cost-efficient way to train models with multiple cheap VMs, trumping both more centralized and powerful hardware and even on-demand cloud offerings at competitive prices.
title How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
Performance
I.2.11; C.2.4; C.4; D.2.8
url https://arxiv.org/abs/2306.03163