Improving training time and GPU utilization in geo-distributed language model training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Palak, Reddy, Tella Rajashekhar, Kataria, Bhaskar, Gandhi, Rohan, Tandon, Karan, Bhattacherjee, Debopam, Padmanabhan, Venkata N.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917023021268992
author Palak
Reddy, Tella Rajashekhar
Kataria, Bhaskar
Gandhi, Rohan
Tandon, Karan
Bhattacherjee, Debopam
Padmanabhan, Venkata N.
author_facet Palak
Reddy, Tella Rajashekhar
Kataria, Bhaskar
Gandhi, Rohan
Tandon, Karan
Bhattacherjee, Debopam
Padmanabhan, Venkata N.
contents The widespread adoption of language models (LMs) has caused a huge surge in demand for GPUs. Training large LMs requires tens of thousands of GPUs and housing them in the same datacenter (DC) is a challenge due to many constraints including availability of peak power. We focus on training such models across multiple DCs connected via the Wide-Area-Network (WAN). We built Atlas that speeds up the training time using novel workload-aware temporal bandwidth sharing and other design choices. While Atlas improves the training time, it does not completely eliminate the bubbles (idle GPU cycles). We built BubbleTea that runs prefill-as-a-service (part of LM inference) during the bubbles thus improving the GPU utilization without any impact on training. Compared to state-of-the-art designs, Atlas and BubbleTea together achieve up to 17x faster training, and up to 94% GPU utilization. The code will be open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14458
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving training time and GPU utilization in geo-distributed language model training
Palak
Reddy, Tella Rajashekhar
Kataria, Bhaskar
Gandhi, Rohan
Tandon, Karan
Bhattacherjee, Debopam
Padmanabhan, Venkata N.
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
The widespread adoption of language models (LMs) has caused a huge surge in demand for GPUs. Training large LMs requires tens of thousands of GPUs and housing them in the same datacenter (DC) is a challenge due to many constraints including availability of peak power. We focus on training such models across multiple DCs connected via the Wide-Area-Network (WAN). We built Atlas that speeds up the training time using novel workload-aware temporal bandwidth sharing and other design choices. While Atlas improves the training time, it does not completely eliminate the bubbles (idle GPU cycles). We built BubbleTea that runs prefill-as-a-service (part of LM inference) during the bubbles thus improving the GPU utilization without any impact on training. Compared to state-of-the-art designs, Atlas and BubbleTea together achieve up to 17x faster training, and up to 94% GPU utilization. The code will be open-sourced.
title Improving training time and GPU utilization in geo-distributed language model training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.14458