Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines"

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grafberger, Stefan, Groth, Paul, Schelter, Sebastian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909185615069184
author Grafberger, Stefan
Groth, Paul
Schelter, Sebastian
author_facet Grafberger, Stefan
Groth, Paul
Schelter, Sebastian
contents Data scientists develop ML pipelines in an iterative manner: they repeatedly screen a pipeline for potential issues, debug it, and then revise and improve its code according to their findings. However, this manual process is tedious and error-prone. Therefore, we propose to support data scientists during this development cycle with automatically derived interactive suggestions for pipeline improvements. We discuss our vision to generate these suggestions with so-called shadow pipelines, hidden variants of the original pipeline that modify it to auto-detect potential issues, try out modifications for improvements, and suggest and explain these modifications to the user. We envision to apply incremental view maintenance-based optimisations to ensure low-latency computation and maintenance of the shadow pipelines. We conduct preliminary experiments to showcase the feasibility of our envisioned approach and the potential benefits of our proposed optimisations.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19591
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines"
Grafberger, Stefan
Groth, Paul
Schelter, Sebastian
Databases
Machine Learning
Software Engineering
H.2; H.2.8; H.4; D.2.6; I.2
Data scientists develop ML pipelines in an iterative manner: they repeatedly screen a pipeline for potential issues, debug it, and then revise and improve its code according to their findings. However, this manual process is tedious and error-prone. Therefore, we propose to support data scientists during this development cycle with automatically derived interactive suggestions for pipeline improvements. We discuss our vision to generate these suggestions with so-called shadow pipelines, hidden variants of the original pipeline that modify it to auto-detect potential issues, try out modifications for improvements, and suggest and explain these modifications to the user. We envision to apply incremental view maintenance-based optimisations to ensure low-latency computation and maintenance of the shadow pipelines. We conduct preliminary experiments to showcase the feasibility of our envisioned approach and the potential benefits of our proposed optimisations.
title Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines"
topic Databases
Machine Learning
Software Engineering
H.2; H.2.8; H.4; D.2.6; I.2
url https://arxiv.org/abs/2404.19591