Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Eungyeup, Sun, Mingjie, Baek, Christina, Raghunathan, Aditi, Kolter, J. Zico
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912109867040768
author Kim, Eungyeup
Sun, Mingjie
Baek, Christina
Raghunathan, Aditi
Kolter, J. Zico
author_facet Kim, Eungyeup
Sun, Mingjie
Baek, Christina
Raghunathan, Aditi
Kolter, J. Zico
contents Recently, Miller et al. (2021) and Baek et al. (2022) empirically demonstrated strong linear correlations between in-distribution (ID) versus out-of-distribution (OOD) accuracy and agreement. These trends, coined accuracy-on-the-line (ACL) and agreement-on-the-line (AGL), enable OOD model selection and performance estimation without labeled data. However, these phenomena also break for certain shifts, such as CIFAR10-C Gaussian Noise, posing a critical bottleneck. In this paper, we make a key finding that recent test-time adaptation (TTA) methods not only improve OOD performance, but drastically strengthen the ACL and AGL trends in models, even in shifts where models showed very weak correlations before. To analyze this, we revisit the theoretical conditions from Miller et al. (2021) that outline the types of distribution shifts needed for perfect ACL in linear models. Surprisingly, these conditions are satisfied after applying TTA to deep models in the penultimate feature embedding space. In particular, TTA causes the data distribution to collapse complex shifts into those can be expressed by a singular scaling variable in the feature space. Our results show that by combining TTA with AGL-based estimation methods, we can estimate the OOD performance of models with high precision for a broader set of distribution shifts. This lends us a simple system for selecting the best hyperparameters and adaptation strategy without any OOD labeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2310_04941
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
Kim, Eungyeup
Sun, Mingjie
Baek, Christina
Raghunathan, Aditi
Kolter, J. Zico
Machine Learning
Artificial Intelligence
Recently, Miller et al. (2021) and Baek et al. (2022) empirically demonstrated strong linear correlations between in-distribution (ID) versus out-of-distribution (OOD) accuracy and agreement. These trends, coined accuracy-on-the-line (ACL) and agreement-on-the-line (AGL), enable OOD model selection and performance estimation without labeled data. However, these phenomena also break for certain shifts, such as CIFAR10-C Gaussian Noise, posing a critical bottleneck. In this paper, we make a key finding that recent test-time adaptation (TTA) methods not only improve OOD performance, but drastically strengthen the ACL and AGL trends in models, even in shifts where models showed very weak correlations before. To analyze this, we revisit the theoretical conditions from Miller et al. (2021) that outline the types of distribution shifts needed for perfect ACL in linear models. Surprisingly, these conditions are satisfied after applying TTA to deep models in the penultimate feature embedding space. In particular, TTA causes the data distribution to collapse complex shifts into those can be expressed by a singular scaling variable in the feature space. Our results show that by combining TTA with AGL-based estimation methods, we can estimate the OOD performance of models with high precision for a broader set of distribution shifts. This lends us a simple system for selecting the best hyperparameters and adaptation strategy without any OOD labeled data.
title Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2310.04941