Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915515589459968 |
|---|---|
| author | Strati, Foteini Zhang, Zhendong Manos, George Périz, Ixeia Sánchez Hu, Qinghao Chen, Tiancheng Buzcu, Berk Han, Song Delgado, Pamela Klimovic, Ana |
| author_facet | Strati, Foteini Zhang, Zhendong Manos, George Périz, Ixeia Sánchez Hu, Qinghao Chen, Tiancheng Buzcu, Berk Han, Song Delgado, Pamela Klimovic, Ana |
| contents | The high GPU demand of ML training makes it hard to allocate large homogeneous clusters of high-end GPUs in a single availability zone. Leveraging heterogeneous GPUs available within and across zones can improve throughput at a reasonable cost. However, training ML models on heterogeneous resources introduces significant challenges, such as stragglers and a large search space of possible job configurations. Current systems lack support for efficiently training models on heterogeneous resources. We present Sailor, a system that automates distributed training over heterogeneous, geo-distributed, and dynamically available resources. Sailor combines an efficient search space exploration algorithm, accurate runtime and memory footprint simulation, and a distributed training framework that supports different types of heterogeneity to optimize training throughput and cost. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_17096 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters Strati, Foteini Zhang, Zhendong Manos, George Périz, Ixeia Sánchez Hu, Qinghao Chen, Tiancheng Buzcu, Berk Han, Song Delgado, Pamela Klimovic, Ana Distributed, Parallel, and Cluster Computing The high GPU demand of ML training makes it hard to allocate large homogeneous clusters of high-end GPUs in a single availability zone. Leveraging heterogeneous GPUs available within and across zones can improve throughput at a reasonable cost. However, training ML models on heterogeneous resources introduces significant challenges, such as stragglers and a large search space of possible job configurations. Current systems lack support for efficiently training models on heterogeneous resources. We present Sailor, a system that automates distributed training over heterogeneous, geo-distributed, and dynamically available resources. Sailor combines an efficient search space exploration algorithm, accurate runtime and memory footprint simulation, and a distributed training framework that supports different types of heterogeneity to optimize training throughput and cost. |
| title | Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2504.17096 |