Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chhabra, Anshuman, Li, Bo, Chen, Jian, Mohapatra, Prasant, Liu, Hongfu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908623457746944
author Chhabra, Anshuman
Li, Bo
Chen, Jian
Mohapatra, Prasant
Liu, Hongfu
author_facet Chhabra, Anshuman
Li, Bo
Chen, Jian
Mohapatra, Prasant
Liu, Hongfu
contents A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for assessing training data influence on model predictions. Despite their widespread use, their high computational cost associated with calculating the inverse of the Hessian matrix pose constraints, particularly when analyzing large-sized deep models. In this paper, we establish a bridge between identifying detrimental training samples via influence functions and outlier gradient detection. This transformation not only presents a straightforward and Hessian-free formulation but also provides insights into the role of the gradient in sample impact. Through systematic empirical evaluations, we first validate the hypothesis of our proposed outlier gradient analysis approach on synthetic datasets. We then demonstrate its effectiveness in detecting mislabeled samples in vision models and selecting data samples for improving performance of natural language processing transformer models. We also extend its use to influential sample identification for fine-tuning Large Language Models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03869
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
Chhabra, Anshuman
Li, Bo
Chen, Jian
Mohapatra, Prasant
Liu, Hongfu
Machine Learning
Artificial Intelligence
A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for assessing training data influence on model predictions. Despite their widespread use, their high computational cost associated with calculating the inverse of the Hessian matrix pose constraints, particularly when analyzing large-sized deep models. In this paper, we establish a bridge between identifying detrimental training samples via influence functions and outlier gradient detection. This transformation not only presents a straightforward and Hessian-free formulation but also provides insights into the role of the gradient in sample impact. Through systematic empirical evaluations, we first validate the hypothesis of our proposed outlier gradient analysis approach on synthetic datasets. We then demonstrate its effectiveness in detecting mislabeled samples in vision models and selecting data samples for improving performance of natural language processing transformer models. We also extend its use to influential sample identification for fine-tuning Large Language Models.
title Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.03869