FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Di Profio, Mattia, Zhong, Mingjun, Sripada, Yaji, Jaspars, Marcel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912511672975360
author Di Profio, Mattia
Zhong, Mingjun
Sripada, Yaji
Jaspars, Marcel
author_facet Di Profio, Mattia
Zhong, Mingjun
Sripada, Yaji
Jaspars, Marcel
contents The Extract, Transform, Load (ETL) workflow is fundamental for populating and maintaining data warehouses and other data stores accessed by analysts for downstream tasks. A major shortcoming of modern ETL solutions is the extensive need for a human-in-the-loop, required to design and implement context-specific, and often non-generalisable transformations. While related work in the field of ETL automation shows promising progress, there is a lack of solutions capable of automatically designing and applying these transformations. We present FlowETL, a novel example-based autonomous ETL pipeline architecture designed to automatically standardise and prepare input datasets according to a concise, user-defined target dataset. FlowETL is an ecosystem of components which interact together to achieve the desired outcome. A Planning Engine uses a paired input-output datasets sample to construct a transformation plan, which is then applied by an ETL worker to the source dataset. Monitoring and logging provide observability throughout the entire pipeline. The results show promising generalisation capabilities across 14 datasets of various domains, file structures, and file sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering
Di Profio, Mattia
Zhong, Mingjun
Sripada, Yaji
Jaspars, Marcel
Software Engineering
The Extract, Transform, Load (ETL) workflow is fundamental for populating and maintaining data warehouses and other data stores accessed by analysts for downstream tasks. A major shortcoming of modern ETL solutions is the extensive need for a human-in-the-loop, required to design and implement context-specific, and often non-generalisable transformations. While related work in the field of ETL automation shows promising progress, there is a lack of solutions capable of automatically designing and applying these transformations. We present FlowETL, a novel example-based autonomous ETL pipeline architecture designed to automatically standardise and prepare input datasets according to a concise, user-defined target dataset. FlowETL is an ecosystem of components which interact together to achieve the desired outcome. A Planning Engine uses a paired input-output datasets sample to construct a transformation plan, which is then applied by an ETL worker to the source dataset. Monitoring and logging provide observability throughout the entire pipeline. The results show promising generalisation capabilities across 14 datasets of various domains, file structures, and file sizes.
title FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering
topic Software Engineering
url https://arxiv.org/abs/2507.23118