A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Loi, Chung Ming, Reinarz, Anne, Lykkegaard, Mikkel, Hornsby, William, Buchanan, James, Seelinger, Linus
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915222112960512
author Loi, Chung Ming
Reinarz, Anne
Lykkegaard, Mikkel
Hornsby, William
Buchanan, James
Seelinger, Linus
author_facet Loi, Chung Ming
Reinarz, Anne
Lykkegaard, Mikkel
Hornsby, William
Buchanan, James
Seelinger, Linus
contents Uncertainty Quantification (UQ) workloads are becoming increasingly common in science and engineering. They involve the submission of thousands or even millions of similar tasks with potentially unpredictable runtimes, where the total number is usually not known a priori. A static one-size-fits-all batch script would likely lead to suboptimal scheduling, and native schedulers installed on High Performance Computing (HPC) systems such as SLURM often struggle to efficiently handle such workloads. In this paper, we introduce a new load balancing approach suitable for UQ workflows. To demonstrate its efficiency in a real-world setting, we focus on the GS2 gyrokinetic plasma turbulence simulator. Individual simulations can be computationally demanding, with runtimes varying significantly-from minutes to hours-depending on the high-dimensional input parameters. Our approach uses UQ and Modelling Bridge, which offers a language-agnostic interface to a simulation model, combined with HyperQueue which works alongside the native scheduler. In particular, deploying this framework on HPC systems does not require system-level changes. We benchmark our proposed framework against a standalone SLURM approach using GS2 and a Gaussian Process surrogate thereof. Our results demonstrate a reduction in scheduling overhead by up to three orders of magnitude and a maximum reduction of 38% in CPU time for long-running simulations compared to the naive SLURM approach, while making no assumptions about the job submission patterns inherent to UQ workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22645
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
Loi, Chung Ming
Reinarz, Anne
Lykkegaard, Mikkel
Hornsby, William
Buchanan, James
Seelinger, Linus
Distributed, Parallel, and Cluster Computing
Uncertainty Quantification (UQ) workloads are becoming increasingly common in science and engineering. They involve the submission of thousands or even millions of similar tasks with potentially unpredictable runtimes, where the total number is usually not known a priori. A static one-size-fits-all batch script would likely lead to suboptimal scheduling, and native schedulers installed on High Performance Computing (HPC) systems such as SLURM often struggle to efficiently handle such workloads. In this paper, we introduce a new load balancing approach suitable for UQ workflows. To demonstrate its efficiency in a real-world setting, we focus on the GS2 gyrokinetic plasma turbulence simulator. Individual simulations can be computationally demanding, with runtimes varying significantly-from minutes to hours-depending on the high-dimensional input parameters. Our approach uses UQ and Modelling Bridge, which offers a language-agnostic interface to a simulation model, combined with HyperQueue which works alongside the native scheduler. In particular, deploying this framework on HPC systems does not require system-level changes. We benchmark our proposed framework against a standalone SLURM approach using GS2 and a Gaussian Process surrogate thereof. Our results demonstrate a reduction in scheduling overhead by up to three orders of magnitude and a maximum reduction of 38% in CPU time for long-running simulations compared to the naive SLURM approach, while making no assumptions about the job submission patterns inherent to UQ workflows.
title A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.22645