The Entropy Enigma: Success and Failure of Entropy Minimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Press, Ori, Shwartz-Ziv, Ravid, LeCun, Yann, Bethge, Matthias
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913347742466048
author Press, Ori
Shwartz-Ziv, Ravid
LeCun, Yann
Bethge, Matthias
author_facet Press, Ori
Shwartz-Ziv, Ravid
LeCun, Yann
Bethge, Matthias
contents Entropy minimization (EM) is frequently used to increase the accuracy of classification models when they're faced with new data at test time. EM is a self-supervised learning method that optimizes classifiers to assign even higher probabilities to their top predicted classes. In this paper, we analyze why EM works when adapting a model for a few steps and why it eventually fails after adapting for many steps. We show that, at first, EM causes the model to embed test images close to training images, thereby increasing model accuracy. After many steps of optimization, EM makes the model embed test images far away from the embeddings of training images, which results in a degradation of accuracy. Building upon our insights, we present a method for solving a practical problem: estimating a model's accuracy on a given arbitrary dataset without having access to its labels. Our method estimates accuracy by looking at how the embeddings of input images change as the model is optimized to minimize entropy. Experiments on 23 challenging datasets show that our method sets the SoTA with a mean absolute error of $5.75\%$, an improvement of $29.62\%$ over the previous SoTA on this task. Our code is available at https://github.com/oripress/EntropyEnigma
format Preprint
id arxiv_https___arxiv_org_abs_2405_05012
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Entropy Enigma: Success and Failure of Entropy Minimization
Press, Ori
Shwartz-Ziv, Ravid
LeCun, Yann
Bethge, Matthias
Computer Vision and Pattern Recognition
Entropy minimization (EM) is frequently used to increase the accuracy of classification models when they're faced with new data at test time. EM is a self-supervised learning method that optimizes classifiers to assign even higher probabilities to their top predicted classes. In this paper, we analyze why EM works when adapting a model for a few steps and why it eventually fails after adapting for many steps. We show that, at first, EM causes the model to embed test images close to training images, thereby increasing model accuracy. After many steps of optimization, EM makes the model embed test images far away from the embeddings of training images, which results in a degradation of accuracy. Building upon our insights, we present a method for solving a practical problem: estimating a model's accuracy on a given arbitrary dataset without having access to its labels. Our method estimates accuracy by looking at how the embeddings of input images change as the model is optimized to minimize entropy. Experiments on 23 challenging datasets show that our method sets the SoTA with a mean absolute error of $5.75\%$, an improvement of $29.62\%$ over the previous SoTA on this task. Our code is available at https://github.com/oripress/EntropyEnigma
title The Entropy Enigma: Success and Failure of Entropy Minimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.05012