DUPRE: Data Utility Prediction for Efficient Data Valuation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pham, Kieu Thao Nguyen, Sim, Rachael Hwee Ling, Nguyen, Quoc Phong, Ng, See Kiong, Low, Bryan Kian Hsiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915375508094976
author Pham, Kieu Thao Nguyen
Sim, Rachael Hwee Ling
Nguyen, Quoc Phong
Ng, See Kiong
Low, Bryan Kian Hsiang
author_facet Pham, Kieu Thao Nguyen
Sim, Rachael Hwee Ling
Nguyen, Quoc Phong
Ng, See Kiong
Low, Bryan Kian Hsiang
contents Data valuation is increasingly used in machine learning (ML) to decide the fair compensation for data owners and identify valuable or harmful data for improving ML models. Cooperative game theory-based data valuation, such as Data Shapley, requires evaluating the data utility (e.g., validation accuracy) and retraining the ML model for multiple data subsets. While most existing works on efficient estimation of the Shapley values have focused on reducing the number of subsets to evaluate, our framework, \texttt{DUPRE}, takes an alternative yet complementary approach that reduces the cost per subset evaluation by predicting data utilities instead of evaluating them by model retraining. Specifically, given the evaluated data utilities of some data subsets, \texttt{DUPRE} fits a \emph{Gaussian process} (GP) regression model to predict the utility of every other data subset. Our key contribution lies in the design of our GP kernel based on the sliced Wasserstein distance between empirical data distributions. In particular, we show that the kernel is valid and positive semi-definite, encodes prior knowledge of similarities between different data subsets, and can be efficiently computed. We empirically verify that \texttt{DUPRE} introduces low prediction error and speeds up data valuation for various ML models, datasets, and utility functions.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16152
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DUPRE: Data Utility Prediction for Efficient Data Valuation
Pham, Kieu Thao Nguyen
Sim, Rachael Hwee Ling
Nguyen, Quoc Phong
Ng, See Kiong
Low, Bryan Kian Hsiang
Machine Learning
Computer Science and Game Theory
Data valuation is increasingly used in machine learning (ML) to decide the fair compensation for data owners and identify valuable or harmful data for improving ML models. Cooperative game theory-based data valuation, such as Data Shapley, requires evaluating the data utility (e.g., validation accuracy) and retraining the ML model for multiple data subsets. While most existing works on efficient estimation of the Shapley values have focused on reducing the number of subsets to evaluate, our framework, \texttt{DUPRE}, takes an alternative yet complementary approach that reduces the cost per subset evaluation by predicting data utilities instead of evaluating them by model retraining. Specifically, given the evaluated data utilities of some data subsets, \texttt{DUPRE} fits a \emph{Gaussian process} (GP) regression model to predict the utility of every other data subset. Our key contribution lies in the design of our GP kernel based on the sliced Wasserstein distance between empirical data distributions. In particular, we show that the kernel is valid and positive semi-definite, encodes prior knowledge of similarities between different data subsets, and can be efficiently computed. We empirically verify that \texttt{DUPRE} introduces low prediction error and speeds up data valuation for various ML models, datasets, and utility functions.
title DUPRE: Data Utility Prediction for Efficient Data Valuation
topic Machine Learning
Computer Science and Game Theory
url https://arxiv.org/abs/2502.16152