Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Vuong, Tung-Long, Phan, Hoang, Vo, Vy, Bui, Anh, Do, Thanh-Toan, Le, Trung, Phung, Dinh
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912427546771456
author Vuong, Tung-Long
Phan, Hoang
Vo, Vy
Bui, Anh
Do, Thanh-Toan
Le, Trung
Phung, Dinh
author_facet Vuong, Tung-Long
Phan, Hoang
Vo, Vy
Bui, Anh
Do, Thanh-Toan
Le, Trung
Phung, Dinh
contents Recent approaches leveraging multi-modal pre-trained models like CLIP for Unsupervised Domain Adaptation (UDA) have shown significant promise in bridging domain gaps and improving generalization by utilizing rich semantic knowledge and robust visual representations learned through extensive pre-training on diverse image-text datasets. While these methods achieve state-of-the-art performance across benchmarks, much of the improvement stems from base pseudo-labels (CLIP zero-shot predictions) and self-training mechanisms. Thus, the training mechanism exhibits a key limitation wherein the visual embedding distribution in target domains can deviate from the visual embedding distribution in the pre-trained model, leading to misguided signals from class descriptions. This work introduces a fresh solution to reinforce these pseudo-labels and facilitate target-prompt learning, by exploiting the geometry of visual and text embeddings - an aspect that is overlooked by existing methods. We first propose to directly leverage the reference predictions (from source prompts) based on the relationship between source and target visual embeddings. We later show that there is a strong clustering behavior observed between visual and text embeddings in pre-trained multi-modal models. Building on optimal transport theory, we transform this insight into a novel strategy to enforce the clustering property in text embeddings, further enhancing the alignment in the target domain. Our experiments and ablation studies validate the effectiveness of the proposed approach, demonstrating superior performance and improved quality of target prompts in terms of representation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation
Vuong, Tung-Long
Phan, Hoang
Vo, Vy
Bui, Anh
Do, Thanh-Toan
Le, Trung
Phung, Dinh
Computer Vision and Pattern Recognition
Recent approaches leveraging multi-modal pre-trained models like CLIP for Unsupervised Domain Adaptation (UDA) have shown significant promise in bridging domain gaps and improving generalization by utilizing rich semantic knowledge and robust visual representations learned through extensive pre-training on diverse image-text datasets. While these methods achieve state-of-the-art performance across benchmarks, much of the improvement stems from base pseudo-labels (CLIP zero-shot predictions) and self-training mechanisms. Thus, the training mechanism exhibits a key limitation wherein the visual embedding distribution in target domains can deviate from the visual embedding distribution in the pre-trained model, leading to misguided signals from class descriptions. This work introduces a fresh solution to reinforce these pseudo-labels and facilitate target-prompt learning, by exploiting the geometry of visual and text embeddings - an aspect that is overlooked by existing methods. We first propose to directly leverage the reference predictions (from source prompts) based on the relationship between source and target visual embeddings. We later show that there is a strong clustering behavior observed between visual and text embeddings in pre-trained multi-modal models. Building on optimal transport theory, we transform this insight into a novel strategy to enforce the clustering property in text embeddings, further enhancing the alignment in the target domain. Our experiments and ablation studies validate the effectiveness of the proposed approach, demonstrating superior performance and improved quality of target prompts in terms of representation.
title Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.11493