Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Xiaoran, Yang, Jiangang, Chong, Wenyue, Shi, Wenhui, Sun, Shichu, Xing, Jing, Liu, Jian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912349383819264
author Xu, Xiaoran
Yang, Jiangang
Chong, Wenyue
Shi, Wenhui
Sun, Shichu
Xing, Jing
Liu, Jian
author_facet Xu, Xiaoran
Yang, Jiangang
Chong, Wenyue
Shi, Wenhui
Sun, Shichu
Xing, Jing
Liu, Jian
contents Single-Domain Generalized Object Detection~(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse unseen target domains, making it suitable for multimedia applications that involve various domain shifts, such as intelligent video surveillance and VR/AR technologies. With the success of large-scale Vision-Language Models, recent S-DGOD approaches exploit pre-trained vision-language knowledge to guide invariant feature learning across visual domains. However, the utilized knowledge remains at a coarse-grained level~(e.g., the textual description of adverse weather paired with the image) and serves as an implicit regularization for guidance, struggling to learn accurate region- and object-level features in varying domains. In this work, we propose a new cross-modal feature learning method, which can capture generalized and discriminative regional features for S-DGOD tasks. The core of our method is the mechanism of Cross-modal and Region-aware Feature Interaction, which simultaneously learns both inter-modal and intra-modal regional invariance through dynamic interactions between fine-grained textual and visual features. Moreover, we design a simple but effective strategy called Cross-domain Proposal Refining and Mixing, which aligns the position of region proposals across multiple domains and diversifies them, enhancing the localization ability of detectors in unseen scenarios. Our method achieves new state-of-the-art results on S-DGOD benchmark datasets, with improvements of +8.8\%~mPC on Cityscapes-C and +7.9\%~mPC on DWD over baselines, demonstrating its efficacy.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19086
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
Xu, Xiaoran
Yang, Jiangang
Chong, Wenyue
Shi, Wenhui
Sun, Shichu
Xing, Jing
Liu, Jian
Computer Vision and Pattern Recognition
Single-Domain Generalized Object Detection~(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse unseen target domains, making it suitable for multimedia applications that involve various domain shifts, such as intelligent video surveillance and VR/AR technologies. With the success of large-scale Vision-Language Models, recent S-DGOD approaches exploit pre-trained vision-language knowledge to guide invariant feature learning across visual domains. However, the utilized knowledge remains at a coarse-grained level~(e.g., the textual description of adverse weather paired with the image) and serves as an implicit regularization for guidance, struggling to learn accurate region- and object-level features in varying domains. In this work, we propose a new cross-modal feature learning method, which can capture generalized and discriminative regional features for S-DGOD tasks. The core of our method is the mechanism of Cross-modal and Region-aware Feature Interaction, which simultaneously learns both inter-modal and intra-modal regional invariance through dynamic interactions between fine-grained textual and visual features. Moreover, we design a simple but effective strategy called Cross-domain Proposal Refining and Mixing, which aligns the position of region proposals across multiple domains and diversifies them, enhancing the localization ability of detectors in unseen scenarios. Our method achieves new state-of-the-art results on S-DGOD benchmark datasets, with improvements of +8.8\%~mPC on Cityscapes-C and +7.9\%~mPC on DWD over baselines, demonstrating its efficacy.
title Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.19086