Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hao, Wang, Jindong, Shah, Ankit, Tao, Ran, Wei, Hongxin, Xie, Xing, Sugiyama, Masashi, Raj, Bhiksha
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929269222932480
author Chen, Hao
Wang, Jindong
Shah, Ankit
Tao, Ran
Wei, Hongxin
Xie, Xing
Sugiyama, Masashi
Raj, Bhiksha
author_facet Chen, Hao
Wang, Jindong
Shah, Ankit
Tao, Ran
Wei, Hongxin
Xie, Xing
Sugiyama, Masashi
Raj, Bhiksha
contents Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-training datasets and to mitigate its impact on downstream tasks. More specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) transfer performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing data distribution are different. We empirically verify that the reason behind is noise in pre-training shapes the feature space differently. We then propose a light-weight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering one may not be able to fully fine-tune or even access the pre-trained models. We conduct practical experiments on popular vision and language models that are pre-trained on noisy data for evaluation of our approach. Our analysis and results show the importance of this interesting and novel research direction, which we term Noisy Model Learning.
format Preprint
id arxiv_https___arxiv_org_abs_2309_17002
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
Chen, Hao
Wang, Jindong
Shah, Ankit
Tao, Ran
Wei, Hongxin
Xie, Xing
Sugiyama, Masashi
Raj, Bhiksha
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-training datasets and to mitigate its impact on downstream tasks. More specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) transfer performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing data distribution are different. We empirically verify that the reason behind is noise in pre-training shapes the feature space differently. We then propose a light-weight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering one may not be able to fully fine-tune or even access the pre-trained models. We conduct practical experiments on popular vision and language models that are pre-trained on noisy data for evaluation of our approach. Our analysis and results show the importance of this interesting and novel research direction, which we term Noisy Model Learning.
title Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.17002