Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908623457746944 |
|---|---|
| author | Chhabra, Anshuman Li, Bo Chen, Jian Mohapatra, Prasant Liu, Hongfu |
| author_facet | Chhabra, Anshuman Li, Bo Chen, Jian Mohapatra, Prasant Liu, Hongfu |
| contents | A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for assessing training data influence on model predictions. Despite their widespread use, their high computational cost associated with calculating the inverse of the Hessian matrix pose constraints, particularly when analyzing large-sized deep models. In this paper, we establish a bridge between identifying detrimental training samples via influence functions and outlier gradient detection. This transformation not only presents a straightforward and Hessian-free formulation but also provides insights into the role of the gradient in sample impact. Through systematic empirical evaluations, we first validate the hypothesis of our proposed outlier gradient analysis approach on synthetic datasets. We then demonstrate its effectiveness in detecting mislabeled samples in vision models and selecting data samples for improving performance of natural language processing transformer models. We also extend its use to influential sample identification for fine-tuning Large Language Models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_03869 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models Chhabra, Anshuman Li, Bo Chen, Jian Mohapatra, Prasant Liu, Hongfu Machine Learning Artificial Intelligence A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for assessing training data influence on model predictions. Despite their widespread use, their high computational cost associated with calculating the inverse of the Hessian matrix pose constraints, particularly when analyzing large-sized deep models. In this paper, we establish a bridge between identifying detrimental training samples via influence functions and outlier gradient detection. This transformation not only presents a straightforward and Hessian-free formulation but also provides insights into the role of the gradient in sample impact. Through systematic empirical evaluations, we first validate the hypothesis of our proposed outlier gradient analysis approach on synthetic datasets. We then demonstrate its effectiveness in detecting mislabeled samples in vision models and selecting data samples for improving performance of natural language processing transformer models. We also extend its use to influential sample identification for fine-tuning Large Language Models. |
| title | Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2405.03869 |