Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Namba, Hiroyuki, Horiguchi, Shota, Hamamoto, Masaki, Egi, Masashi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916123319992320
author Namba, Hiroyuki
Horiguchi, Shota
Hamamoto, Masaki
Egi, Masashi
author_facet Namba, Hiroyuki
Horiguchi, Shota
Hamamoto, Masaki
Egi, Masashi
contents Data cleansing aims to improve model performance by removing a set of harmful instances from the training dataset. Data Shapley is a common theoretically guaranteed method to evaluate the contribution of each instance to model performance; however, it requires training on all subsets of the training data, which is computationally expensive. In this paper, we propose an iterativemethod to fast identify a subset of instances with low data Shapley values by using the thresholding bandit algorithm. We provide a theoretical guarantee that the proposed method can accurately select harmful instances if a sufficiently large number of iterations is conducted. Empirical evaluation using various models and datasets demonstrated that the proposed method efficiently improved the computational speed while maintaining the model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08209
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits
Namba, Hiroyuki
Horiguchi, Shota
Hamamoto, Masaki
Egi, Masashi
Machine Learning
Artificial Intelligence
Data cleansing aims to improve model performance by removing a set of harmful instances from the training dataset. Data Shapley is a common theoretically guaranteed method to evaluate the contribution of each instance to model performance; however, it requires training on all subsets of the training data, which is computationally expensive. In this paper, we propose an iterativemethod to fast identify a subset of instances with low data Shapley values by using the thresholding bandit algorithm. We provide a theoretical guarantee that the proposed method can accurately select harmful instances if a sufficiently large number of iterations is conducted. Empirical evaluation using various models and datasets demonstrated that the proposed method efficiently improved the computational speed while maintaining the model performance.
title Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.08209