UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Xiao, Wei, Fei, Wang, Yong, Zhao, Wenda, Li, Feiyi, Chu, Xiangxiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916853613330432
author Zhang, Xiao
Wei, Fei
Wang, Yong
Zhao, Wenda
Li, Feiyi
Chu, Xiangxiang
author_facet Zhang, Xiao
Wei, Fei
Wang, Yong
Zhao, Wenda
Li, Feiyi
Chu, Xiangxiang
contents Zero-shot domain adaptation (ZSDA) presents substantial challenges due to the lack of images in the target domain. Previous approaches leverage Vision-Language Models (VLMs) to tackle this challenge, exploiting their zero-shot learning capabilities. However, these methods primarily address domain distribution shifts and overlook the misalignment between the detection task and VLMs, which rely on manually crafted prompts. To overcome these limitations, we propose the unified prompt and representation enhancement (UPRE) framework, which jointly optimizes both textual prompts and visual representations. Specifically, our approach introduces a multi-view domain prompt that combines linguistic domain priors with detection-specific knowledge, and a visual representation enhancement module that produces domain style variations. Furthermore, we introduce multi-level enhancement strategies, including relative domain distance and positive-negative separation, which align multi-modal representations at the image level and capture diverse visual representations at the instance level, respectively. Extensive experiments conducted on nine benchmark datasets demonstrate the superior performance of our framework in ZSDA detection scenarios. Code is available at https://github.com/AMAP-ML/UPRE.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
Zhang, Xiao
Wei, Fei
Wang, Yong
Zhao, Wenda
Li, Feiyi
Chu, Xiangxiang
Computer Vision and Pattern Recognition
Zero-shot domain adaptation (ZSDA) presents substantial challenges due to the lack of images in the target domain. Previous approaches leverage Vision-Language Models (VLMs) to tackle this challenge, exploiting their zero-shot learning capabilities. However, these methods primarily address domain distribution shifts and overlook the misalignment between the detection task and VLMs, which rely on manually crafted prompts. To overcome these limitations, we propose the unified prompt and representation enhancement (UPRE) framework, which jointly optimizes both textual prompts and visual representations. Specifically, our approach introduces a multi-view domain prompt that combines linguistic domain priors with detection-specific knowledge, and a visual representation enhancement module that produces domain style variations. Furthermore, we introduce multi-level enhancement strategies, including relative domain distance and positive-negative separation, which align multi-modal representations at the image level and capture diverse visual representations at the instance level, respectively. Extensive experiments conducted on nine benchmark datasets demonstrate the superior performance of our framework in ZSDA detection scenarios. Code is available at https://github.com/AMAP-ML/UPRE.
title UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.00721