Automated Planning for Optimal Data Pipeline Instantiation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Amado, Leonardo Rosa, Vogel, Adriano, Griebler, Dalvan, Licks, Gabriel Paludo, Simon, Eric, Meneguzzi, Felipe
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914246291357696
author Amado, Leonardo Rosa
Vogel, Adriano
Griebler, Dalvan
Licks, Gabriel Paludo
Simon, Eric
Meneguzzi, Felipe
author_facet Amado, Leonardo Rosa
Vogel, Adriano
Griebler, Dalvan
Licks, Gabriel Paludo
Simon, Eric
Meneguzzi, Felipe
contents Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing resources to be allocated in a data center, ideally minimizing the overhead for communicating data and executing operators in the pipeline while considering each operator's execution requirements. In this paper, we model the problem of optimal data pipeline deployment as planning with action costs, where we propose heuristics aiming to minimize total execution time. Experimental results indicate that the heuristics can outperform the baseline deployment and that a heuristic based on connections outperforms other strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Planning for Optimal Data Pipeline Instantiation
Amado, Leonardo Rosa
Vogel, Adriano
Griebler, Dalvan
Licks, Gabriel Paludo
Simon, Eric
Meneguzzi, Felipe
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing resources to be allocated in a data center, ideally minimizing the overhead for communicating data and executing operators in the pipeline while considering each operator's execution requirements. In this paper, we model the problem of optimal data pipeline deployment as planning with action costs, where we propose heuristics aiming to minimize total execution time. Experimental results indicate that the heuristics can outperform the baseline deployment and that a heuristic based on connections outperforms other strategies.
title Automated Planning for Optimal Data Pipeline Instantiation
topic Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.12626