GPU Cluster Scheduling for Network-Sensitive Deep Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912694278291456 |
|---|---|
| author | Sharma, Aakash Bhasi, Vivek M. Singh, Sonali Kesidis, George Kandemir, Mahmut T. Das, Chita R. |
| author_facet | Sharma, Aakash Bhasi, Vivek M. Singh, Sonali Kesidis, George Kandemir, Mahmut T. Das, Chita R. |
| contents | We propose a novel GPU-cluster scheduler for distributed DL (DDL) workloads that enables proximity based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication-network delays. Our scheduler consists of three major components: (i) a classical delay scheduling algorithm to facilitate job placement and consolidation; (ii) a network-sensitive job preemption strategy; and (iii) an "auto-tuner" mechanism to optimize delay timers for effective delay scheduling. Additionally, to enable a cost-effective methodology for large-scale experiments, we develop a data-driven DDL cluster simulation platform. Employing the simulation platform we compare against several state-of-the-art alternatives on real-world workload traces to demonstrate the benefits of our design. Our scheduler can provide improvement of up to 69% in end-to-end Makespan for training all jobs compared to the prevailing consolidation-based scheduling methods, while reducing the average job completion time by up to 83% and minimizing the communication overheads by up to 98% under congested networking conditions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_16492 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GPU Cluster Scheduling for Network-Sensitive Deep Learning Sharma, Aakash Bhasi, Vivek M. Singh, Sonali Kesidis, George Kandemir, Mahmut T. Das, Chita R. Performance Distributed, Parallel, and Cluster Computing Machine Learning We propose a novel GPU-cluster scheduler for distributed DL (DDL) workloads that enables proximity based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication-network delays. Our scheduler consists of three major components: (i) a classical delay scheduling algorithm to facilitate job placement and consolidation; (ii) a network-sensitive job preemption strategy; and (iii) an "auto-tuner" mechanism to optimize delay timers for effective delay scheduling. Additionally, to enable a cost-effective methodology for large-scale experiments, we develop a data-driven DDL cluster simulation platform. Employing the simulation platform we compare against several state-of-the-art alternatives on real-world workload traces to demonstrate the benefits of our design. Our scheduler can provide improvement of up to 69% in end-to-end Makespan for training all jobs compared to the prevailing consolidation-based scheduling methods, while reducing the average job completion time by up to 83% and minimizing the communication overheads by up to 98% under congested networking conditions. |
| title | GPU Cluster Scheduling for Network-Sensitive Deep Learning |
| topic | Performance Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2401.16492 |