Zero-Execution Retrieval-Augmented Configuration Tuning of Spark Applications

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Suri, Raunaq, Gofman, Ilan, Yu, Guangwei, Cresswell, Jesse C.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917946619592704
author Suri, Raunaq
Gofman, Ilan
Yu, Guangwei
Cresswell, Jesse C.
author_facet Suri, Raunaq
Gofman, Ilan
Yu, Guangwei
Cresswell, Jesse C.
contents Large-scale data processing is increasingly done using distributed computing frameworks like Apache Spark, which have a considerable number of configurable parameters that affect runtime performance. For optimal performance, these parameters must be tuned to the specific job being run. Tuning commonly requires multiple executions to collect runtime information for updating parameters. This is infeasible for ad hoc queries that are run once or infrequently. Zero-execution tuning, where parameters are automatically set before a job's first run, can provide significant savings for all types of applications, but is more challenging since runtime information is not available. In this work, we propose a novel method for zero-execution tuning of Spark configurations based on retrieval. Our method achieves 93.3% of the runtime improvement of state-of-the-art one-execution optimization, entirely avoiding the slow initial execution using default settings. The shift to zero-execution tuning results in a lower cumulative runtime over the first 140 runs, and provides the largest benefit for ad hoc and analytical queries which only need to be executed once. We release the largest and most comprehensive suite of Spark query datasets, optimal configurations, and runtime information, which will promote future development of zero-execution tuning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03826
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Execution Retrieval-Augmented Configuration Tuning of Spark Applications
Suri, Raunaq
Gofman, Ilan
Yu, Guangwei
Cresswell, Jesse C.
Distributed, Parallel, and Cluster Computing
Large-scale data processing is increasingly done using distributed computing frameworks like Apache Spark, which have a considerable number of configurable parameters that affect runtime performance. For optimal performance, these parameters must be tuned to the specific job being run. Tuning commonly requires multiple executions to collect runtime information for updating parameters. This is infeasible for ad hoc queries that are run once or infrequently. Zero-execution tuning, where parameters are automatically set before a job's first run, can provide significant savings for all types of applications, but is more challenging since runtime information is not available. In this work, we propose a novel method for zero-execution tuning of Spark configurations based on retrieval. Our method achieves 93.3% of the runtime improvement of state-of-the-art one-execution optimization, entirely avoiding the slow initial execution using default settings. The shift to zero-execution tuning results in a lower cumulative runtime over the first 140 runs, and provides the largest benefit for ad hoc and analytical queries which only need to be executed once. We release the largest and most comprehensive suite of Spark query datasets, optimal configurations, and runtime information, which will promote future development of zero-execution tuning methods.
title Zero-Execution Retrieval-Augmented Configuration Tuning of Spark Applications
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.03826