MRSch: Multi-Resource Scheduling for HPC

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Boyang, Fan, Yuping, Dearing, Matthew, Lan, Zhiling, Richy, Paul, Allcocky, William, Papka, Michael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913298228707328
author Li, Boyang
Fan, Yuping
Dearing, Matthew
Lan, Zhiling
Richy, Paul
Allcocky, William
Papka, Michael
author_facet Li, Boyang
Fan, Yuping
Dearing, Matthew
Lan, Zhiling
Richy, Paul
Allcocky, William
Papka, Michael
contents Emerging workloads in high-performance computing (HPC) are embracing significant changes, such as having diverse resource requirements instead of being CPU-centric. This advancement forces cluster schedulers to consider multiple schedulable resources during decision-making. Existing scheduling studies rely on heuristic or optimization methods, which are limited by an inability to adapt to new scenarios for ensuring long-term scheduling performance. We present an intelligent scheduling agent named MRSch for multi-resource scheduling in HPC that leverages direct future prediction (DFP), an advanced multi-objective reinforcement learning algorithm. While DFP demonstrated outstanding performance in a gaming competition, it has not been previously explored in the context of HPC scheduling. Several key techniques are developed in this study to tackle the challenges involved in multi-resource scheduling. These techniques enable MRSch to learn an appropriate scheduling policy automatically and dynamically adapt its policy in response to workload changes via dynamic resource prioritizing. We compare MRSch with existing scheduling methods through extensive tracebase simulations. Our results demonstrate that MRSch improves scheduling performance by up to 48% compared to the existing scheduling methods.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16298
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MRSch: Multi-Resource Scheduling for HPC
Li, Boyang
Fan, Yuping
Dearing, Matthew
Lan, Zhiling
Richy, Paul
Allcocky, William
Papka, Michael
Distributed, Parallel, and Cluster Computing
Emerging workloads in high-performance computing (HPC) are embracing significant changes, such as having diverse resource requirements instead of being CPU-centric. This advancement forces cluster schedulers to consider multiple schedulable resources during decision-making. Existing scheduling studies rely on heuristic or optimization methods, which are limited by an inability to adapt to new scenarios for ensuring long-term scheduling performance. We present an intelligent scheduling agent named MRSch for multi-resource scheduling in HPC that leverages direct future prediction (DFP), an advanced multi-objective reinforcement learning algorithm. While DFP demonstrated outstanding performance in a gaming competition, it has not been previously explored in the context of HPC scheduling. Several key techniques are developed in this study to tackle the challenges involved in multi-resource scheduling. These techniques enable MRSch to learn an appropriate scheduling policy automatically and dynamically adapt its policy in response to workload changes via dynamic resource prioritizing. We compare MRSch with existing scheduling methods through extensive tracebase simulations. Our results demonstrate that MRSch improves scheduling performance by up to 48% compared to the existing scheduling methods.
title MRSch: Multi-Resource Scheduling for HPC
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2403.16298