WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Jiawei, Niu, Yulei, Huang, Shiyuan, Han, Guangxing, Chang, Shih-Fu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910461119692800
author Ma, Jiawei
Niu, Yulei
Huang, Shiyuan
Han, Guangxing
Chang, Shih-Fu
author_facet Ma, Jiawei
Niu, Yulei
Huang, Shiyuan
Han, Guangxing
Chang, Shih-Fu
contents Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mostly at coarse-grained level and ignores visual details, the resulted embeddings are still ineffective in overcoming complexity of domains at inference time. We present a self-supervision framework WIDIn, Wording Images for Domain-Invariant representation, to disentangle discriminative visual representation, by only leveraging data in a single domain and without any test prior. Specifically, for each image, we first estimate the language embedding with fine-grained alignment, which can be consequently used to adaptively identify and then remove domain-specific counterpart from the raw visual embedding. WIDIn can be applied to both pretrained vision-language models like CLIP, and separately trained uni-modal models like MoCo and BERT. Experimental studies on three domain generalization datasets demonstrate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18405
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
Ma, Jiawei
Niu, Yulei
Huang, Shiyuan
Han, Guangxing
Chang, Shih-Fu
Computer Vision and Pattern Recognition
Artificial Intelligence
Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mostly at coarse-grained level and ignores visual details, the resulted embeddings are still ineffective in overcoming complexity of domains at inference time. We present a self-supervision framework WIDIn, Wording Images for Domain-Invariant representation, to disentangle discriminative visual representation, by only leveraging data in a single domain and without any test prior. Specifically, for each image, we first estimate the language embedding with fine-grained alignment, which can be consequently used to adaptively identify and then remove domain-specific counterpart from the raw visual embedding. WIDIn can be applied to both pretrained vision-language models like CLIP, and separately trained uni-modal models like MoCo and BERT. Experimental studies on three domain generalization datasets demonstrate the effectiveness of our approach.
title WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.18405