"What is Different Between These Datasets?" A Framework for Explaining Data Distribution Shifts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Babbar, Varun, Guo, Zhicheng, Rudin, Cynthia
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909800290320384
author Babbar, Varun
Guo, Zhicheng
Rudin, Cynthia
author_facet Babbar, Varun
Guo, Zhicheng
Rudin, Cynthia
contents The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two datasets from the same domain may exhibit differing distributions. While many techniques exist for detecting such distribution shifts, there is a lack of comprehensive methods to explain these differences in a human-understandable way beyond opaque quantitative metrics. To bridge this gap, we propose a versatile framework of interpretable methods for comparing datasets. Using a variety of case studies, we demonstrate the effectiveness of our approach across diverse data modalities-including tabular data, text data, images, time-series signals -- in both low and high-dimensional settings. These methods complement existing techniques by providing actionable and interpretable insights to better understand and address distribution shifts.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05652
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle "What is Different Between These Datasets?" A Framework for Explaining Data Distribution Shifts
Babbar, Varun
Guo, Zhicheng
Rudin, Cynthia
Machine Learning
Artificial Intelligence
The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two datasets from the same domain may exhibit differing distributions. While many techniques exist for detecting such distribution shifts, there is a lack of comprehensive methods to explain these differences in a human-understandable way beyond opaque quantitative metrics. To bridge this gap, we propose a versatile framework of interpretable methods for comparing datasets. Using a variety of case studies, we demonstrate the effectiveness of our approach across diverse data modalities-including tabular data, text data, images, time-series signals -- in both low and high-dimensional settings. These methods complement existing techniques by providing actionable and interpretable insights to better understand and address distribution shifts.
title "What is Different Between These Datasets?" A Framework for Explaining Data Distribution Shifts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2403.05652