Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines"
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909185615069184 |
|---|---|
| author | Grafberger, Stefan Groth, Paul Schelter, Sebastian |
| author_facet | Grafberger, Stefan Groth, Paul Schelter, Sebastian |
| contents | Data scientists develop ML pipelines in an iterative manner: they repeatedly screen a pipeline for potential issues, debug it, and then revise and improve its code according to their findings. However, this manual process is tedious and error-prone. Therefore, we propose to support data scientists during this development cycle with automatically derived interactive suggestions for pipeline improvements. We discuss our vision to generate these suggestions with so-called shadow pipelines, hidden variants of the original pipeline that modify it to auto-detect potential issues, try out modifications for improvements, and suggest and explain these modifications to the user. We envision to apply incremental view maintenance-based optimisations to ensure low-latency computation and maintenance of the shadow pipelines. We conduct preliminary experiments to showcase the feasibility of our envisioned approach and the potential benefits of our proposed optimisations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_19591 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines" Grafberger, Stefan Groth, Paul Schelter, Sebastian Databases Machine Learning Software Engineering H.2; H.2.8; H.4; D.2.6; I.2 Data scientists develop ML pipelines in an iterative manner: they repeatedly screen a pipeline for potential issues, debug it, and then revise and improve its code according to their findings. However, this manual process is tedious and error-prone. Therefore, we propose to support data scientists during this development cycle with automatically derived interactive suggestions for pipeline improvements. We discuss our vision to generate these suggestions with so-called shadow pipelines, hidden variants of the original pipeline that modify it to auto-detect potential issues, try out modifications for improvements, and suggest and explain these modifications to the user. We envision to apply incremental view maintenance-based optimisations to ensure low-latency computation and maintenance of the shadow pipelines. We conduct preliminary experiments to showcase the feasibility of our envisioned approach and the potential benefits of our proposed optimisations. |
| title | Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines" |
| topic | Databases Machine Learning Software Engineering H.2; H.2.8; H.4; D.2.6; I.2 |
| url | https://arxiv.org/abs/2404.19591 |