UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Haoyu, Cheng, Zhi-Qi, Moreira, Gabriel, Zhu, Jiawen, Sun, Jingdong, Ren, Bukun, He, Jun-Yan, Dai, Qi, Hua, Xian-Sheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929630099800064
author Jiang, Haoyu
Cheng, Zhi-Qi
Moreira, Gabriel
Zhu, Jiawen
Sun, Jingdong
Ren, Bukun
He, Jun-Yan
Dai, Qi
Hua, Xian-Sheng
author_facet Jiang, Haoyu
Cheng, Zhi-Qi
Moreira, Gabriel
Zhu, Jiawen
Sun, Jingdong
Ren, Bukun
He, Jun-Yan
Dai, Qi
Hua, Xian-Sheng
contents Universal Cross-Domain Retrieval (UCDR) retrieves relevant images from unseen domains and classes without semantic labels, ensuring robust generalization. Existing methods commonly employ prompt tuning with pre-trained vision-language models but are inherently limited by static prompts, reducing adaptability. We propose UCDR-Adapter, which enhances pre-trained models with adapters and dynamic prompt generation through a two-phase training strategy. First, Source Adapter Learning integrates class semantics with domain-specific visual knowledge using a Learnable Textual Semantic Template and optimizes Class and Domain Prompts via momentum updates and dual loss functions for robust alignment. Second, Target Prompt Generation creates dynamic prompts by attending to masked source prompts, enabling seamless adaptation to unseen domains and classes. Unlike prior approaches, UCDR-Adapter dynamically adapts to evolving data distributions, enhancing both flexibility and generalization. During inference, only the image branch and generated prompts are used, eliminating reliance on textual inputs for highly efficient retrieval. Extensive benchmark experiments show that UCDR-Adapter consistently outperforms ProS in most cases and other state-of-the-art methods on UCDR, U(c)CDR, and U(d)CDR settings.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10680
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
Jiang, Haoyu
Cheng, Zhi-Qi
Moreira, Gabriel
Zhu, Jiawen
Sun, Jingdong
Ren, Bukun
He, Jun-Yan
Dai, Qi
Hua, Xian-Sheng
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
Universal Cross-Domain Retrieval (UCDR) retrieves relevant images from unseen domains and classes without semantic labels, ensuring robust generalization. Existing methods commonly employ prompt tuning with pre-trained vision-language models but are inherently limited by static prompts, reducing adaptability. We propose UCDR-Adapter, which enhances pre-trained models with adapters and dynamic prompt generation through a two-phase training strategy. First, Source Adapter Learning integrates class semantics with domain-specific visual knowledge using a Learnable Textual Semantic Template and optimizes Class and Domain Prompts via momentum updates and dual loss functions for robust alignment. Second, Target Prompt Generation creates dynamic prompts by attending to masked source prompts, enabling seamless adaptation to unseen domains and classes. Unlike prior approaches, UCDR-Adapter dynamically adapts to evolving data distributions, enhancing both flexibility and generalization. During inference, only the image branch and generated prompts are used, eliminating reliance on textual inputs for highly efficient retrieval. Extensive benchmark experiments show that UCDR-Adapter consistently outperforms ProS in most cases and other state-of-the-art methods on UCDR, U(c)CDR, and U(d)CDR settings.
title UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
topic Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
url https://arxiv.org/abs/2412.10680