Attention Bootstrapping for Multi-Modal Test-Time Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yusheng, Luo, Junyu, Luo, Xiao, Huang, Jinsheng, Yuan, Jingyang, Xiao, Zhiping, Zhang, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916642433269760
author Zhao, Yusheng
Luo, Junyu
Luo, Xiao
Huang, Jinsheng
Yuan, Jingyang
Xiao, Zhiping
Zhang, Ming
author_facet Zhao, Yusheng
Luo, Junyu
Luo, Xiao
Huang, Jinsheng
Yuan, Jingyang
Xiao, Zhiping
Zhang, Ming
contents Test-time adaptation aims to adapt a well-trained model to potential distribution shifts at test time using only unlabeled test data, without access to the original training data. While previous efforts mainly focus on a single modality, test-time distribution shift in the multi-modal setting is more complex and calls for new solutions. This paper tackles the problem of multi-modal test-time adaptation by proposing a novel method named Attention Bootstrapping with Principal Entropy Minimization (ABPEM). We observe that test-time distribution shift causes misalignment across modalities, leading to a large gap between intra-modality discrepancies (measured by self-attention) and inter-modality discrepancies (measured by cross-attention). We name this the attention gap. This attention gap widens with more severe distribution shifts, hindering effective modality fusion. To mitigate this attention gap and encourage better modality fusion, we propose attention bootstrapping that promotes cross-attention with the guidance of self-attention. Moreover, to reduce the gradient noise in the commonly-used entropy minimization, we adopt principal entropy minimization, a refinement of entropy minimization that reduces gradient noise by focusing on the principal parts of entropy, excluding less reliable gradient information. Extensive experiments on the benchmarks validate the effectiveness of the proposed ABPEM in comparison with competing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02221
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attention Bootstrapping for Multi-Modal Test-Time Adaptation
Zhao, Yusheng
Luo, Junyu
Luo, Xiao
Huang, Jinsheng
Yuan, Jingyang
Xiao, Zhiping
Zhang, Ming
Artificial Intelligence
Test-time adaptation aims to adapt a well-trained model to potential distribution shifts at test time using only unlabeled test data, without access to the original training data. While previous efforts mainly focus on a single modality, test-time distribution shift in the multi-modal setting is more complex and calls for new solutions. This paper tackles the problem of multi-modal test-time adaptation by proposing a novel method named Attention Bootstrapping with Principal Entropy Minimization (ABPEM). We observe that test-time distribution shift causes misalignment across modalities, leading to a large gap between intra-modality discrepancies (measured by self-attention) and inter-modality discrepancies (measured by cross-attention). We name this the attention gap. This attention gap widens with more severe distribution shifts, hindering effective modality fusion. To mitigate this attention gap and encourage better modality fusion, we propose attention bootstrapping that promotes cross-attention with the guidance of self-attention. Moreover, to reduce the gradient noise in the commonly-used entropy minimization, we adopt principal entropy minimization, a refinement of entropy minimization that reduces gradient noise by focusing on the principal parts of entropy, excluding less reliable gradient information. Extensive experiments on the benchmarks validate the effectiveness of the proposed ABPEM in comparison with competing baselines.
title Attention Bootstrapping for Multi-Modal Test-Time Adaptation
topic Artificial Intelligence
url https://arxiv.org/abs/2503.02221