Rethinking Domain Adaptation and Generalization in the Era of CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Ruoyu, Yu, Tao, Jin, Xin, Yu, Xiaoyuan, Xiao, Lei, Chen, Zhibo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929429599485952
author Feng, Ruoyu
Yu, Tao
Jin, Xin
Yu, Xiaoyuan
Xiao, Lei
Chen, Zhibo
author_facet Feng, Ruoyu
Yu, Tao
Jin, Xin
Yu, Xiaoyuan
Xiao, Lei
Chen, Zhibo
contents In recent studies on domain adaptation, significant emphasis has been placed on the advancement of learning shared knowledge from a source domain to a target domain. Recently, the large vision-language pre-trained model, i.e., CLIP has shown strong ability on zero-shot recognition, and parameter efficient tuning can further improve its performance on specific tasks. This work demonstrates that a simple domain prior boosts CLIP's zero-shot recognition in a specific domain. Besides, CLIP's adaptation relies less on source domain data due to its diverse pre-training dataset. Furthermore, we create a benchmark for zero-shot adaptation and pseudo-labeling based self-training with CLIP. Last but not least, we propose to improve the task generalization ability of CLIP from multiple unlabeled domains, which is a more practical and unique scenario. We believe our findings motivate a rethinking of domain adaptation benchmarks and the associated role of related algorithms in the era of CLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15173
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Domain Adaptation and Generalization in the Era of CLIP
Feng, Ruoyu
Yu, Tao
Jin, Xin
Yu, Xiaoyuan
Xiao, Lei
Chen, Zhibo
Computer Vision and Pattern Recognition
In recent studies on domain adaptation, significant emphasis has been placed on the advancement of learning shared knowledge from a source domain to a target domain. Recently, the large vision-language pre-trained model, i.e., CLIP has shown strong ability on zero-shot recognition, and parameter efficient tuning can further improve its performance on specific tasks. This work demonstrates that a simple domain prior boosts CLIP's zero-shot recognition in a specific domain. Besides, CLIP's adaptation relies less on source domain data due to its diverse pre-training dataset. Furthermore, we create a benchmark for zero-shot adaptation and pseudo-labeling based self-training with CLIP. Last but not least, we propose to improve the task generalization ability of CLIP from multiple unlabeled domains, which is a more practical and unique scenario. We believe our findings motivate a rethinking of domain adaptation benchmarks and the associated role of related algorithms in the era of CLIP.
title Rethinking Domain Adaptation and Generalization in the Era of CLIP
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.15173