Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Strati, Foteini, Zhang, Zhendong, Manos, George, Périz, Ixeia Sánchez, Hu, Qinghao, Chen, Tiancheng, Buzcu, Berk, Han, Song, Delgado, Pamela, Klimovic, Ana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915515589459968
author Strati, Foteini
Zhang, Zhendong
Manos, George
Périz, Ixeia Sánchez
Hu, Qinghao
Chen, Tiancheng
Buzcu, Berk
Han, Song
Delgado, Pamela
Klimovic, Ana
author_facet Strati, Foteini
Zhang, Zhendong
Manos, George
Périz, Ixeia Sánchez
Hu, Qinghao
Chen, Tiancheng
Buzcu, Berk
Han, Song
Delgado, Pamela
Klimovic, Ana
contents The high GPU demand of ML training makes it hard to allocate large homogeneous clusters of high-end GPUs in a single availability zone. Leveraging heterogeneous GPUs available within and across zones can improve throughput at a reasonable cost. However, training ML models on heterogeneous resources introduces significant challenges, such as stragglers and a large search space of possible job configurations. Current systems lack support for efficiently training models on heterogeneous resources. We present Sailor, a system that automates distributed training over heterogeneous, geo-distributed, and dynamically available resources. Sailor combines an efficient search space exploration algorithm, accurate runtime and memory footprint simulation, and a distributed training framework that supports different types of heterogeneity to optimize training throughput and cost.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
Strati, Foteini
Zhang, Zhendong
Manos, George
Périz, Ixeia Sánchez
Hu, Qinghao
Chen, Tiancheng
Buzcu, Berk
Han, Song
Delgado, Pamela
Klimovic, Ana
Distributed, Parallel, and Cluster Computing
The high GPU demand of ML training makes it hard to allocate large homogeneous clusters of high-end GPUs in a single availability zone. Leveraging heterogeneous GPUs available within and across zones can improve throughput at a reasonable cost. However, training ML models on heterogeneous resources introduces significant challenges, such as stragglers and a large search space of possible job configurations. Current systems lack support for efficiently training models on heterogeneous resources. We present Sailor, a system that automates distributed training over heterogeneous, geo-distributed, and dynamically available resources. Sailor combines an efficient search space exploration algorithm, accurate runtime and memory footprint simulation, and a distributed training framework that supports different types of heterogeneity to optimize training throughput and cost.
title Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2504.17096