Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, De, Xu, Zhipeng, Jiang, Xinyang, Li, Dongsheng, Wang, Nannan, Gao, Xinbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908431590359040
author Cheng, De
Xu, Zhipeng
Jiang, Xinyang
Li, Dongsheng
Wang, Nannan
Gao, Xinbo
author_facet Cheng, De
Xu, Zhipeng
Jiang, Xinyang
Li, Dongsheng
Wang, Nannan
Gao, Xinbo
contents Domain Generalization (DG) seeks to develop a versatile model capable of performing effectively on unseen target domains. Notably, recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have demonstrated considerable potential in enhancing the generalization capabilities of deep learning models. Despite the increasing attention toward VFM-based domain prompt tuning within DG, the effective design of prompts capable of disentangling invariant features across diverse domains remains a critical challenge. In this paper, we propose addressing this challenge by leveraging the controllable and flexible language prompt of the VFM. Noting that the text modality of VFMs is naturally easier to disentangle, we introduce a novel framework for text feature-guided visual prompt tuning. This framework first automatically disentangles the text prompt using a large language model (LLM) and then learns domain-invariant visual representation guided by the disentangled text feature. However, relying solely on language to guide visual feature disentanglement has limitations, as visual features can sometimes be too complex or nuanced to be fully captured by descriptive text. To address this, we introduce Worst Explicit Representation Alignment (WERA), which extends text-guided visual prompts by incorporating an additional set of abstract prompts. These prompts enhance source domain diversity through stylized image augmentations, while alignment constraints ensure that visual representations remain consistent across both the original and augmented distributions. Experiments conducted on major DG datasets, including PACS, VLCS, OfficeHome, DomainNet, and TerraInc, demonstrate that our proposed method outperforms state-of-the-art DG methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02288
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
Cheng, De
Xu, Zhipeng
Jiang, Xinyang
Li, Dongsheng
Wang, Nannan
Gao, Xinbo
Computer Vision and Pattern Recognition
Machine Learning
Domain Generalization (DG) seeks to develop a versatile model capable of performing effectively on unseen target domains. Notably, recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have demonstrated considerable potential in enhancing the generalization capabilities of deep learning models. Despite the increasing attention toward VFM-based domain prompt tuning within DG, the effective design of prompts capable of disentangling invariant features across diverse domains remains a critical challenge. In this paper, we propose addressing this challenge by leveraging the controllable and flexible language prompt of the VFM. Noting that the text modality of VFMs is naturally easier to disentangle, we introduce a novel framework for text feature-guided visual prompt tuning. This framework first automatically disentangles the text prompt using a large language model (LLM) and then learns domain-invariant visual representation guided by the disentangled text feature. However, relying solely on language to guide visual feature disentanglement has limitations, as visual features can sometimes be too complex or nuanced to be fully captured by descriptive text. To address this, we introduce Worst Explicit Representation Alignment (WERA), which extends text-guided visual prompts by incorporating an additional set of abstract prompts. These prompts enhance source domain diversity through stylized image augmentations, while alignment constraints ensure that visual representations remain consistent across both the original and augmented distributions. Experiments conducted on major DG datasets, including PACS, VLCS, OfficeHome, DomainNet, and TerraInc, demonstrate that our proposed method outperforms state-of-the-art DG methods.
title Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.02288