Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lehmann, Fabian, Bader, Jonathan, De Mecquenem, Ninon, Wang, Xing, Bountris, Vasilis, Friederici, Florian, Leser, Ulf, Thamsen, Lauritz
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910643361153024
author Lehmann, Fabian
Bader, Jonathan
De Mecquenem, Ninon
Wang, Xing
Bountris, Vasilis
Friederici, Florian
Leser, Ulf
Thamsen, Lauritz
author_facet Lehmann, Fabian
Bader, Jonathan
De Mecquenem, Ninon
Wang, Xing
Bountris, Vasilis
Friederici, Florian
Leser, Ulf
Thamsen, Lauritz
contents Scientific workflows are used to analyze large amounts of data. These workflows comprise numerous tasks, many of which are executed repeatedly, running the same custom program on different inputs. Users specify resource allocations for each task, which must be sufficient for all inputs to prevent task failures. As a result, task memory allocations tend to be overly conservative, wasting precious cluster resources, limiting overall parallelism, and increasing workflow makespan. In this paper, we first benchmark a state-of-the-art method on four real-life workflows from the nf-core workflow repository. This analysis reveals that certain assumptions underlying current prediction methods, which typically were evaluated only on simulated workflows, cannot generally be confirmed for real workflows and executions. We then present Ponder, a new online task-sizing strategy that considers and chooses between different methods to cater to different memory demand patterns. We implemented Ponder for Nextflow and made the code publicly available. In an experimental evaluation that also considers the impact of memory predictions on scheduling, Ponder improves Memory Allocation Quality on average by 71.0% and makespan by 21.8% in comparison to a state-of-the-art method. Moreover, Ponder produces 93.8% fewer task failures.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00047
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
Lehmann, Fabian
Bader, Jonathan
De Mecquenem, Ninon
Wang, Xing
Bountris, Vasilis
Friederici, Florian
Leser, Ulf
Thamsen, Lauritz
Distributed, Parallel, and Cluster Computing
Scientific workflows are used to analyze large amounts of data. These workflows comprise numerous tasks, many of which are executed repeatedly, running the same custom program on different inputs. Users specify resource allocations for each task, which must be sufficient for all inputs to prevent task failures. As a result, task memory allocations tend to be overly conservative, wasting precious cluster resources, limiting overall parallelism, and increasing workflow makespan. In this paper, we first benchmark a state-of-the-art method on four real-life workflows from the nf-core workflow repository. This analysis reveals that certain assumptions underlying current prediction methods, which typically were evaluated only on simulated workflows, cannot generally be confirmed for real workflows and executions. We then present Ponder, a new online task-sizing strategy that considers and chooses between different methods to cater to different memory demand patterns. We implemented Ponder for Nextflow and made the code publicly available. In an experimental evaluation that also considers the impact of memory predictions on scheduling, Ponder improves Memory Allocation Quality on average by 71.0% and makespan by 21.8% in comparison to a state-of-the-art method. Moreover, Ponder produces 93.8% fewer task failures.
title Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2408.00047