Generalizability of experimental studies

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Matteucci, Federico, Arzamasov, Vadim, Cribeiro-Ramallo, Jose, Heyden, Marco, Ntounas, Konstantin, Böhm, Klemens
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918230524690432
author Matteucci, Federico
Arzamasov, Vadim
Cribeiro-Ramallo, Jose
Heyden, Marco
Ntounas, Konstantin
Böhm, Klemens
author_facet Matteucci, Federico
Arzamasov, Vadim
Cribeiro-Ramallo, Jose
Heyden, Marco
Ntounas, Konstantin
Böhm, Klemens
contents Experimental studies are a cornerstone of Machine Learning (ML) research. A common and often implicit assumption is that the study's results will generalize beyond the study itself, e.g., to new data. That is, repeating the same study under different conditions will likely yield similar results. Existing frameworks to measure generalizability, borrowed from the casual inference literature, cannot capture the complexity of the results and the goals of an ML study. The problem of measuring generalizability in the more general ML setting is thus still open, also due to the lack of a mathematical formalization of experimental studies. In this paper, we propose such a formalization, use it to develop a framework to quantify generalizability, and propose an instantiation based on rankings and the Maximum Mean Discrepancy. We show how our framework offers insights into the number of experiments necessary for a generalizable study, and how experimenters can benefit from it. Finally, we release the genexpy Python package, which allows for an effortless evaluation of the generalizability of other experimental studies.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17374
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generalizability of experimental studies
Matteucci, Federico
Arzamasov, Vadim
Cribeiro-Ramallo, Jose
Heyden, Marco
Ntounas, Konstantin
Böhm, Klemens
Machine Learning
Statistics Theory
Experimental studies are a cornerstone of Machine Learning (ML) research. A common and often implicit assumption is that the study's results will generalize beyond the study itself, e.g., to new data. That is, repeating the same study under different conditions will likely yield similar results. Existing frameworks to measure generalizability, borrowed from the casual inference literature, cannot capture the complexity of the results and the goals of an ML study. The problem of measuring generalizability in the more general ML setting is thus still open, also due to the lack of a mathematical formalization of experimental studies. In this paper, we propose such a formalization, use it to develop a framework to quantify generalizability, and propose an instantiation based on rankings and the Maximum Mean Discrepancy. We show how our framework offers insights into the number of experiments necessary for a generalizable study, and how experimenters can benefit from it. Finally, we release the genexpy Python package, which allows for an effortless evaluation of the generalizability of other experimental studies.
title Generalizability of experimental studies
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2406.17374