GPU Cluster Scheduling for Network-Sensitive Deep Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharma, Aakash, Bhasi, Vivek M., Singh, Sonali, Kesidis, George, Kandemir, Mahmut T., Das, Chita R.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912694278291456
author Sharma, Aakash
Bhasi, Vivek M.
Singh, Sonali
Kesidis, George
Kandemir, Mahmut T.
Das, Chita R.
author_facet Sharma, Aakash
Bhasi, Vivek M.
Singh, Sonali
Kesidis, George
Kandemir, Mahmut T.
Das, Chita R.
contents We propose a novel GPU-cluster scheduler for distributed DL (DDL) workloads that enables proximity based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication-network delays. Our scheduler consists of three major components: (i) a classical delay scheduling algorithm to facilitate job placement and consolidation; (ii) a network-sensitive job preemption strategy; and (iii) an "auto-tuner" mechanism to optimize delay timers for effective delay scheduling. Additionally, to enable a cost-effective methodology for large-scale experiments, we develop a data-driven DDL cluster simulation platform. Employing the simulation platform we compare against several state-of-the-art alternatives on real-world workload traces to demonstrate the benefits of our design. Our scheduler can provide improvement of up to 69% in end-to-end Makespan for training all jobs compared to the prevailing consolidation-based scheduling methods, while reducing the average job completion time by up to 83% and minimizing the communication overheads by up to 98% under congested networking conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16492
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GPU Cluster Scheduling for Network-Sensitive Deep Learning
Sharma, Aakash
Bhasi, Vivek M.
Singh, Sonali
Kesidis, George
Kandemir, Mahmut T.
Das, Chita R.
Performance
Distributed, Parallel, and Cluster Computing
Machine Learning
We propose a novel GPU-cluster scheduler for distributed DL (DDL) workloads that enables proximity based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication-network delays. Our scheduler consists of three major components: (i) a classical delay scheduling algorithm to facilitate job placement and consolidation; (ii) a network-sensitive job preemption strategy; and (iii) an "auto-tuner" mechanism to optimize delay timers for effective delay scheduling. Additionally, to enable a cost-effective methodology for large-scale experiments, we develop a data-driven DDL cluster simulation platform. Employing the simulation platform we compare against several state-of-the-art alternatives on real-world workload traces to demonstrate the benefits of our design. Our scheduler can provide improvement of up to 69% in end-to-end Makespan for training all jobs compared to the prevailing consolidation-based scheduling methods, while reducing the average job completion time by up to 83% and minimizing the communication overheads by up to 98% under congested networking conditions.
title GPU Cluster Scheduling for Network-Sensitive Deep Learning
topic Performance
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2401.16492