IT$^3$: Idempotent Test-Time Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Durasov, Nikita, Shocher, Assaf, Oner, Doruk, Chechik, Gal, Efros, Alexei A., Fua, Pascal
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916755941621760
author Durasov, Nikita
Shocher, Assaf
Oner, Doruk
Chechik, Gal
Efros, Alexei A.
Fua, Pascal
author_facet Durasov, Nikita
Shocher, Assaf
Oner, Doruk
Chechik, Gal
Efros, Alexei A.
Fua, Pascal
contents Deep learning models often struggle when deployed in real-world settings due to distribution shifts between training and test data. While existing approaches like domain adaptation and test-time training (TTT) offer partial solutions, they typically require additional data or domain-specific auxiliary tasks. We present Idempotent Test-Time Training (IT$^3$), a novel approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design. Our key insight is that enforcing idempotence -- where repeated applications of a function yield the same result -- can effectively replace domain-specific auxiliary tasks used in previous TTT methods. We theoretically connect idempotence to prediction confidence and demonstrate that minimizing the distance between successive applications of our model during inference leads to improved out-of-distribution performance. Extensive experiments across diverse domains (including image classification, aerodynamics prediction, and aerial segmentation) and architectures (MLPs, CNNs, GNNs) show that IT$^3$ consistently outperforms existing approaches while being simpler and more widely applicable. Our results suggest that idempotence provides a universal principle for test-time adaptation that generalizes across domains and architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04201
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IT$^3$: Idempotent Test-Time Training
Durasov, Nikita
Shocher, Assaf
Oner, Doruk
Chechik, Gal
Efros, Alexei A.
Fua, Pascal
Computer Vision and Pattern Recognition
Deep learning models often struggle when deployed in real-world settings due to distribution shifts between training and test data. While existing approaches like domain adaptation and test-time training (TTT) offer partial solutions, they typically require additional data or domain-specific auxiliary tasks. We present Idempotent Test-Time Training (IT$^3$), a novel approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design. Our key insight is that enforcing idempotence -- where repeated applications of a function yield the same result -- can effectively replace domain-specific auxiliary tasks used in previous TTT methods. We theoretically connect idempotence to prediction confidence and demonstrate that minimizing the distance between successive applications of our model during inference leads to improved out-of-distribution performance. Extensive experiments across diverse domains (including image classification, aerodynamics prediction, and aerial segmentation) and architectures (MLPs, CNNs, GNNs) show that IT$^3$ consistently outperforms existing approaches while being simpler and more widely applicable. Our results suggest that idempotence provides a universal principle for test-time adaptation that generalizes across domains and architectures.
title IT$^3$: Idempotent Test-Time Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.04201