Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Renchunzi, Odonnat, Ambroise, Feofanov, Vasilii, Redko, Ievgen, Zhang, Jianfeng, An, Bo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918016232456192
author Xie, Renchunzi
Odonnat, Ambroise
Feofanov, Vasilii
Redko, Ievgen
Zhang, Jianfeng
An, Bo
author_facet Xie, Renchunzi
Odonnat, Ambroise
Feofanov, Vasilii
Redko, Ievgen
Zhang, Jianfeng
An, Bo
contents Estimating the test performance of a model, possibly under distribution shift, without having access to the ground-truth labels is a challenging, yet very important problem for the safe deployment of machine learning algorithms in the wild. Existing works mostly rely on information from either the outputs or the extracted features of neural networks to estimate a score that correlates with the ground-truth test accuracy. In this paper, we investigate -- both empirically and theoretically -- how the information provided by the gradients can be predictive of the ground-truth test accuracy even under distribution shifts. More specifically, we use the norm of classification-layer gradients, backpropagated from the cross-entropy loss after only one gradient step over test data. Our intuition is that these gradients should be of higher magnitude when the model generalizes poorly. We provide the theoretical insights behind our approach and the key ingredients that ensure its empirical success. Extensive experiments conducted with various architectures on diverse distribution shifts demonstrate that our method significantly outperforms current state-of-the-art approaches. The code is available at https://github.com/Renchunzi-Xie/GdScore
format Preprint
id arxiv_https___arxiv_org_abs_2401_08909
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift
Xie, Renchunzi
Odonnat, Ambroise
Feofanov, Vasilii
Redko, Ievgen
Zhang, Jianfeng
An, Bo
Machine Learning
Estimating the test performance of a model, possibly under distribution shift, without having access to the ground-truth labels is a challenging, yet very important problem for the safe deployment of machine learning algorithms in the wild. Existing works mostly rely on information from either the outputs or the extracted features of neural networks to estimate a score that correlates with the ground-truth test accuracy. In this paper, we investigate -- both empirically and theoretically -- how the information provided by the gradients can be predictive of the ground-truth test accuracy even under distribution shifts. More specifically, we use the norm of classification-layer gradients, backpropagated from the cross-entropy loss after only one gradient step over test data. Our intuition is that these gradients should be of higher magnitude when the model generalizes poorly. We provide the theoretical insights behind our approach and the key ingredients that ensure its empirical success. Extensive experiments conducted with various architectures on diverse distribution shifts demonstrate that our method significantly outperforms current state-of-the-art approaches. The code is available at https://github.com/Renchunzi-Xie/GdScore
title Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift
topic Machine Learning
url https://arxiv.org/abs/2401.08909