Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Namba, Hiroyuki, Horiguchi, Shota, Hamamoto, Masaki, Egi, Masashi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916123319992320
author Namba, Hiroyuki
Horiguchi, Shota
Hamamoto, Masaki
Egi, Masashi
author_facet Namba, Hiroyuki
Horiguchi, Shota
Hamamoto, Masaki
Egi, Masashi
contents Data cleansing aims to improve model performance by removing a set of harmful instances from the training dataset. Data Shapley is a common theoretically guaranteed method to evaluate the contribution of each instance to model performance; however, it requires training on all subsets of the training data, which is computationally expensive. In this paper, we propose an iterativemethod to fast identify a subset of instances with low data Shapley values by using the thresholding bandit algorithm. We provide a theoretical guarantee that the proposed method can accurately select harmful instances if a sufficiently large number of iterations is conducted. Empirical evaluation using various models and datasets demonstrated that the proposed method efficiently improved the computational speed while maintaining the model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08209
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits
Namba, Hiroyuki
Horiguchi, Shota
Hamamoto, Masaki
Egi, Masashi
Machine Learning
Artificial Intelligence
Data cleansing aims to improve model performance by removing a set of harmful instances from the training dataset. Data Shapley is a common theoretically guaranteed method to evaluate the contribution of each instance to model performance; however, it requires training on all subsets of the training data, which is computationally expensive. In this paper, we propose an iterativemethod to fast identify a subset of instances with low data Shapley values by using the thresholding bandit algorithm. We provide a theoretical guarantee that the proposed method can accurately select harmful instances if a sufficiently large number of iterations is conducted. Empirical evaluation using various models and datasets demonstrated that the proposed method efficiently improved the computational speed while maintaining the model performance.
title Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.08209