Subspace Alignment for Vision-Language Model Test-time Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Zhichen, Bao, Wenxuan, Lin, Xiao, Qiu, Ruizhong, Wei, Tianxin, Ning, Xuying, Yan, Yuchen, Luo, Chen, Cheng, Monica Xiao, He, Jingrui, Tong, Hanghang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914250978492416
author Zeng, Zhichen
Bao, Wenxuan
Lin, Xiao
Qiu, Ruizhong
Wei, Tianxin
Ning, Xuying
Yan, Yuchen
Luo, Chen
Cheng, Monica Xiao
He, Jingrui
Tong, Hanghang
author_facet Zeng, Zhichen
Bao, Wenxuan
Lin, Xiao
Qiu, Ruizhong
Wei, Tianxin
Ning, Xuying
Yan, Yuchen
Luo, Chen
Cheng, Monica Xiao
He, Jingrui
Tong, Hanghang
contents Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptation (TTA) emerges as a predominant strategy to adapt VLMs to unlabeled test data on the fly. However, existing TTA methods heavily rely on zero-shot predictions as pseudo-labels for self-training, which can be unreliable under distribution shifts and misguide adaptation due to two fundamental limitations. First (Modality Gap), distribution shifts induce gaps between visual and textual modalities, making cross-modal relations inaccurate. Second (Visual Nuisance), visual embeddings encode rich but task-irrelevant noise that often overwhelms task-specific semantics under distribution shifts. To address these limitations, we propose SubTTA, which aligns the semantic subspaces of both modalities to enhance zero-shot predictions to better guide the TTA process. To bridge the modality gap, SubTTA extracts the principal subspaces of both modalities and aligns the visual manifold to the textual semantic anchor by minimizing their chordal distance. To eliminate visual nuisance, SubTTA projects the aligned visual features onto the task-specific textual subspace, which filters out task-irrelevant noise by constraining visual embeddings within the valid semantic span, and standard TTA is further performed on the purified space to refine the decision boundaries. Extensive experiments on various benchmarks and VLM architectures demonstrate the effectiveness of SubTTA, yielding an average improvement of 2.24% over state-of-the-art TTA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08139
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Subspace Alignment for Vision-Language Model Test-time Adaptation
Zeng, Zhichen
Bao, Wenxuan
Lin, Xiao
Qiu, Ruizhong
Wei, Tianxin
Ning, Xuying
Yan, Yuchen
Luo, Chen
Cheng, Monica Xiao
He, Jingrui
Tong, Hanghang
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptation (TTA) emerges as a predominant strategy to adapt VLMs to unlabeled test data on the fly. However, existing TTA methods heavily rely on zero-shot predictions as pseudo-labels for self-training, which can be unreliable under distribution shifts and misguide adaptation due to two fundamental limitations. First (Modality Gap), distribution shifts induce gaps between visual and textual modalities, making cross-modal relations inaccurate. Second (Visual Nuisance), visual embeddings encode rich but task-irrelevant noise that often overwhelms task-specific semantics under distribution shifts. To address these limitations, we propose SubTTA, which aligns the semantic subspaces of both modalities to enhance zero-shot predictions to better guide the TTA process. To bridge the modality gap, SubTTA extracts the principal subspaces of both modalities and aligns the visual manifold to the textual semantic anchor by minimizing their chordal distance. To eliminate visual nuisance, SubTTA projects the aligned visual features onto the task-specific textual subspace, which filters out task-irrelevant noise by constraining visual embeddings within the valid semantic span, and standard TTA is further performed on the purified space to refine the decision boundaries. Extensive experiments on various benchmarks and VLM architectures demonstrate the effectiveness of SubTTA, yielding an average improvement of 2.24% over state-of-the-art TTA methods.
title Subspace Alignment for Vision-Language Model Test-time Adaptation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.08139