On Leakage in Machine Learning Pipelines

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sasse, Leonard, Nicolaisen-Sobesky, Eliana, Dukart, Juergen, Eickhoff, Simon B., Götz, Michael, Hamdan, Sami, Komeyer, Vera, Kulkarni, Abhijit, Lahnakoski, Juha, Love, Bradley C., Raimondo, Federico, Patil, Kaustubh R.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911096580866048
author Sasse, Leonard
Nicolaisen-Sobesky, Eliana
Dukart, Juergen
Eickhoff, Simon B.
Götz, Michael
Hamdan, Sami
Komeyer, Vera
Kulkarni, Abhijit
Lahnakoski, Juha
Love, Bradley C.
Raimondo, Federico
Patil, Kaustubh R.
author_facet Sasse, Leonard
Nicolaisen-Sobesky, Eliana
Dukart, Juergen
Eickhoff, Simon B.
Götz, Michael
Hamdan, Sami
Komeyer, Vera
Kulkarni, Abhijit
Lahnakoski, Juha
Love, Bradley C.
Raimondo, Federico
Patil, Kaustubh R.
contents Machine learning (ML) provides powerful tools for predictive modeling. ML's popularity stems from the promise of sample-level prediction with applications across a variety of fields from physics and marketing to healthcare. However, if not properly implemented and evaluated, ML pipelines may contain leakage typically resulting in overoptimistic performance estimates and failure to generalize to new data. This can have severe negative financial and societal implications. Our aim is to expand understanding associated with causes leading to leakage when designing, implementing, and evaluating ML pipelines. Illustrated by concrete examples, we provide a comprehensive overview and discussion of various types of leakage that may arise in ML pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2311_04179
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On Leakage in Machine Learning Pipelines
Sasse, Leonard
Nicolaisen-Sobesky, Eliana
Dukart, Juergen
Eickhoff, Simon B.
Götz, Michael
Hamdan, Sami
Komeyer, Vera
Kulkarni, Abhijit
Lahnakoski, Juha
Love, Bradley C.
Raimondo, Federico
Patil, Kaustubh R.
Machine Learning
Artificial Intelligence
Machine learning (ML) provides powerful tools for predictive modeling. ML's popularity stems from the promise of sample-level prediction with applications across a variety of fields from physics and marketing to healthcare. However, if not properly implemented and evaluated, ML pipelines may contain leakage typically resulting in overoptimistic performance estimates and failure to generalize to new data. This can have severe negative financial and societal implications. Our aim is to expand understanding associated with causes leading to leakage when designing, implementing, and evaluating ML pipelines. Illustrated by concrete examples, we provide a comprehensive overview and discussion of various types of leakage that may arise in ML pipelines.
title On Leakage in Machine Learning Pipelines
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2311.04179