Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Will, Jonathan, Treide, Nico, Thamsen, Lauritz, Kao, Odej
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913664589627392
author Will, Jonathan
Treide, Nico
Thamsen, Lauritz
Kao, Odej
author_facet Will, Jonathan
Treide, Nico
Thamsen, Lauritz
Kao, Odej
contents Distributed dataflow systems like Spark and Flink enable data-parallel processing of large datasets on clusters. Yet, selecting appropriate computational resources for dataflow jobs is often challenging. For efficient execution, individual resource allocations, such as memory and CPU cores, must meet the specific resource requirements of the job. An alternative to selecting a static resource allocation for a job execution is autoscaling as implemented for example by Spark. In this paper, we evaluate the resource efficiency of autoscaling batch data processing jobs based on resource demand both conceptually and experimentally by analyzing a new dataset of Spark job executions on Google Dataproc Serverless. In our experimental evaluation, we show that there is no significant resource efficiency gain over static resource allocations. We found that the inherent conceptual limitations of such autoscaling approaches are the inelasticity of node size as well as the inelasticity of the ratio of memory to CPU cores.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
Will, Jonathan
Treide, Nico
Thamsen, Lauritz
Kao, Odej
Distributed, Parallel, and Cluster Computing
C.2.4; I.2.8; I.2.6
Distributed dataflow systems like Spark and Flink enable data-parallel processing of large datasets on clusters. Yet, selecting appropriate computational resources for dataflow jobs is often challenging. For efficient execution, individual resource allocations, such as memory and CPU cores, must meet the specific resource requirements of the job. An alternative to selecting a static resource allocation for a job execution is autoscaling as implemented for example by Spark. In this paper, we evaluate the resource efficiency of autoscaling batch data processing jobs based on resource demand both conceptually and experimentally by analyzing a new dataset of Spark job executions on Google Dataproc Serverless. In our experimental evaluation, we show that there is no significant resource efficiency gain over static resource allocations. We found that the inherent conceptual limitations of such autoscaling approaches are the inelasticity of node size as well as the inelasticity of the ratio of memory to CPU cores.
title Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
topic Distributed, Parallel, and Cluster Computing
C.2.4; I.2.8; I.2.6
url https://arxiv.org/abs/2501.14456