An Empirical Evaluation of Serverless Cloud Infrastructure for Large-Scale Data Processing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bodner, Thomas, Radig, Theo, Justen, David, Ritter, Daniel, Rabl, Tilmann
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912187590639616
author Bodner, Thomas
Radig, Theo
Justen, David
Ritter, Daniel
Rabl, Tilmann
author_facet Bodner, Thomas
Radig, Theo
Justen, David
Ritter, Daniel
Rabl, Tilmann
contents Data processing systems are increasingly deployed in the cloud. While monolithic systems run fully on virtual servers, recent systems embrace cloud infrastructure and utilize the disaggregation of compute and storage to scale them independently. The introduction of serverless compute services, such as AWS Lambda, enables finer-grained and elastic scalability within these systems. Prior work shows the viability of serverless infrastructure for scalable data processing yet also sees limitations due to variable performance and cost overhead, in particular for networking and storage. In this paper, we perform a detailed analysis of the performance and cost characteristics of serverless infrastructure in the data processing context. We base our analysis on a large series of micro-benchmarks across different compute and storage services, as well as end-to-end workloads. To enable our analysis, we propose the Skyrise serverless evaluation platform. For the widely used serverless infrastructure of AWS, our analysis reveals distinct boundaries for performance variability in serverless networks and storage. We further present cost break-even points for serverless compute and storage. These insights provide guidance on when and how serverless infrastructure can be efficiently used for data processing.
format Preprint
id arxiv_https___arxiv_org_abs_2501_07771
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Evaluation of Serverless Cloud Infrastructure for Large-Scale Data Processing
Bodner, Thomas
Radig, Theo
Justen, David
Ritter, Daniel
Rabl, Tilmann
Databases
Data processing systems are increasingly deployed in the cloud. While monolithic systems run fully on virtual servers, recent systems embrace cloud infrastructure and utilize the disaggregation of compute and storage to scale them independently. The introduction of serverless compute services, such as AWS Lambda, enables finer-grained and elastic scalability within these systems. Prior work shows the viability of serverless infrastructure for scalable data processing yet also sees limitations due to variable performance and cost overhead, in particular for networking and storage. In this paper, we perform a detailed analysis of the performance and cost characteristics of serverless infrastructure in the data processing context. We base our analysis on a large series of micro-benchmarks across different compute and storage services, as well as end-to-end workloads. To enable our analysis, we propose the Skyrise serverless evaluation platform. For the widely used serverless infrastructure of AWS, our analysis reveals distinct boundaries for performance variability in serverless networks and storage. We further present cost break-even points for serverless compute and storage. These insights provide guidance on when and how serverless infrastructure can be efficiently used for data processing.
title An Empirical Evaluation of Serverless Cloud Infrastructure for Large-Scale Data Processing
topic Databases
url https://arxiv.org/abs/2501.07771