Automated Planning for Optimal Data Pipeline Instantiation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914246291357696 |
|---|---|
| author | Amado, Leonardo Rosa Vogel, Adriano Griebler, Dalvan Licks, Gabriel Paludo Simon, Eric Meneguzzi, Felipe |
| author_facet | Amado, Leonardo Rosa Vogel, Adriano Griebler, Dalvan Licks, Gabriel Paludo Simon, Eric Meneguzzi, Felipe |
| contents | Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing resources to be allocated in a data center, ideally minimizing the overhead for communicating data and executing operators in the pipeline while considering each operator's execution requirements. In this paper, we model the problem of optimal data pipeline deployment as planning with action costs, where we propose heuristics aiming to minimize total execution time. Experimental results indicate that the heuristics can outperform the baseline deployment and that a heuristic based on connections outperforms other strategies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_12626 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Automated Planning for Optimal Data Pipeline Instantiation Amado, Leonardo Rosa Vogel, Adriano Griebler, Dalvan Licks, Gabriel Paludo Simon, Eric Meneguzzi, Felipe Artificial Intelligence Distributed, Parallel, and Cluster Computing Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing resources to be allocated in a data center, ideally minimizing the overhead for communicating data and executing operators in the pipeline while considering each operator's execution requirements. In this paper, we model the problem of optimal data pipeline deployment as planning with action costs, where we propose heuristics aiming to minimize total execution time. Experimental results indicate that the heuristics can outperform the baseline deployment and that a heuristic based on connections outperforms other strategies. |
| title | Automated Planning for Optimal Data Pipeline Instantiation |
| topic | Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2503.12626 |