Flock: A Low-Cost Streaming Query Engine on FaaS Platforms

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liao, Gang, Deshpande, Amol, Abadi, Daniel J.
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909177088049152
author Liao, Gang
Deshpande, Amol
Abadi, Daniel J.
author_facet Liao, Gang
Deshpande, Amol
Abadi, Daniel J.
contents Existing serverless data analytics systems rely on external storage services like S3 for data shuffling and communication between cloud functions. While this approach provides the elasticity benefits of serverless computing, it incurs additional latency and cost overheads. We present Flock, a novel cloud-native streaming query engine that leverages the on-demand scalability of FaaS platforms for real-time data analytics. Flock utilizes function invocation payloads for efficient data exchange, eliminating the need for external storage. This not only reduces latency and cost but also simplifies the architecture by removing the requirement for a centralized coordinator. Flock employs a template-based approach to dynamically create cloud functions for each query stage and a function group mechanism for handling data aggregation and shuffling. It supports both SQL and DataFrame APIs, making it easy to use. Our evaluation shows that Flock provides significant performance gains and cost savings compared to existing serverless and serverful streaming systems. It outperforms Apache Flink by 10-20x in cost while achieving similar latency and throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2312_16735
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Flock: A Low-Cost Streaming Query Engine on FaaS Platforms
Liao, Gang
Deshpande, Amol
Abadi, Daniel J.
Databases
Distributed, Parallel, and Cluster Computing
Existing serverless data analytics systems rely on external storage services like S3 for data shuffling and communication between cloud functions. While this approach provides the elasticity benefits of serverless computing, it incurs additional latency and cost overheads. We present Flock, a novel cloud-native streaming query engine that leverages the on-demand scalability of FaaS platforms for real-time data analytics. Flock utilizes function invocation payloads for efficient data exchange, eliminating the need for external storage. This not only reduces latency and cost but also simplifies the architecture by removing the requirement for a centralized coordinator. Flock employs a template-based approach to dynamically create cloud functions for each query stage and a function group mechanism for handling data aggregation and shuffling. It supports both SQL and DataFrame APIs, making it easy to use. Our evaluation shows that Flock provides significant performance gains and cost savings compared to existing serverless and serverful streaming systems. It outperforms Apache Flink by 10-20x in cost while achieving similar latency and throughput.
title Flock: A Low-Cost Streaming Query Engine on FaaS Platforms
topic Databases
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2312.16735