Closing the Confusion Loop: CLIP-Guided Alignment for Source-Free Domain Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shanshan, Feng, Ziying, Shen, Xiaozheng, Yang, Xun, Wang, Pichao, He, Zhenwei, Zhang, Xingyi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914315800412160
author Wang, Shanshan
Feng, Ziying
Shen, Xiaozheng
Yang, Xun
Wang, Pichao
He, Zhenwei
Zhang, Xingyi
author_facet Wang, Shanshan
Feng, Ziying
Shen, Xiaozheng
Yang, Xun
Wang, Pichao
He, Zhenwei
Zhang, Xingyi
contents Source-Free Domain Adaptation (SFDA) tackles the problem of adapting a pre-trained source model to an unlabeled target domain without accessing any source data, which is quite suitable for the field of data security. Although recent advances have shown that pseudo-labeling strategies can be effective, they often fail in fine-grained scenarios due to subtle inter-class similarities. A critical but underexplored issue is the presence of asymmetric and dynamic class confusion, where visually similar classes are unequally and inconsistently misclassified by the source model. Existing methods typically ignore such confusion patterns, leading to noisy pseudo-labels and poor target discrimination. To address this, we propose CLIP-Guided Alignment(CGA), a novel framework that explicitly models and mitigates class confusion in SFDA. Generally, our method consists of three parts: (1) MCA: detects first directional confusion pairs by analyzing the predictions of the source model in the target domain; (2) MCC: leverages CLIP to construct confusion-aware textual prompts (e.g. a truck that looks like a bus), enabling more context-sensitive pseudo-labeling; and (3) FAM: builds confusion-guided feature banks for both CLIP and the source model and aligns them using contrastive learning to reduce ambiguity in the representation space. Extensive experiments on various datasets demonstrate that CGA consistently outperforms state-of-the-art SFDA methods, with especially notable gains in confusion-prone and fine-grained scenarios. Our results highlight the importance of explicitly modeling inter-class confusion for effective source-free adaptation. Our code can be find at https://github.com/soloiro/CGA
format Preprint
id arxiv_https___arxiv_org_abs_2602_08730
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Closing the Confusion Loop: CLIP-Guided Alignment for Source-Free Domain Adaptation
Wang, Shanshan
Feng, Ziying
Shen, Xiaozheng
Yang, Xun
Wang, Pichao
He, Zhenwei
Zhang, Xingyi
Computer Vision and Pattern Recognition
Source-Free Domain Adaptation (SFDA) tackles the problem of adapting a pre-trained source model to an unlabeled target domain without accessing any source data, which is quite suitable for the field of data security. Although recent advances have shown that pseudo-labeling strategies can be effective, they often fail in fine-grained scenarios due to subtle inter-class similarities. A critical but underexplored issue is the presence of asymmetric and dynamic class confusion, where visually similar classes are unequally and inconsistently misclassified by the source model. Existing methods typically ignore such confusion patterns, leading to noisy pseudo-labels and poor target discrimination. To address this, we propose CLIP-Guided Alignment(CGA), a novel framework that explicitly models and mitigates class confusion in SFDA. Generally, our method consists of three parts: (1) MCA: detects first directional confusion pairs by analyzing the predictions of the source model in the target domain; (2) MCC: leverages CLIP to construct confusion-aware textual prompts (e.g. a truck that looks like a bus), enabling more context-sensitive pseudo-labeling; and (3) FAM: builds confusion-guided feature banks for both CLIP and the source model and aligns them using contrastive learning to reduce ambiguity in the representation space. Extensive experiments on various datasets demonstrate that CGA consistently outperforms state-of-the-art SFDA methods, with especially notable gains in confusion-prone and fine-grained scenarios. Our results highlight the importance of explicitly modeling inter-class confusion for effective source-free adaptation. Our code can be find at https://github.com/soloiro/CGA
title Closing the Confusion Loop: CLIP-Guided Alignment for Source-Free Domain Adaptation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.08730