BlenDA: Domain Adaptive Object Detection through diffusion-based blending

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tzuhsuan, Huang, Chen-Che, Ku, Chung-Hao, Chen, Jun-Cheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916096438697984
author Huang, Tzuhsuan
Huang, Chen-Che
Ku, Chung-Hao
Chen, Jun-Cheng
author_facet Huang, Tzuhsuan
Huang, Chen-Che
Ku, Chung-Hao
Chen, Jun-Cheng
contents Unsupervised domain adaptation (UDA) aims to transfer a model learned using labeled data from the source domain to unlabeled data in the target domain. To address the large domain gap issue between the source and target domains, we propose a novel regularization method for domain adaptive object detection, BlenDA, by generating the pseudo samples of the intermediate domains and their corresponding soft domain labels for adaptation training. The intermediate samples are generated by dynamically blending the source images with their corresponding translated images using an off-the-shelf pre-trained text-to-image diffusion model which takes the text label of the target domain as input and has demonstrated superior image-to-image translation quality. Based on experimental results from two adaptation benchmarks, our proposed approach can significantly enhance the performance of the state-of-the-art domain adaptive object detector, Adversarial Query Transformer (AQT). Particularly, in the Cityscapes to Foggy Cityscapes adaptation, we achieve an impressive 53.4% mAP on the Foggy Cityscapes dataset, surpassing the previous state-of-the-art by 1.5%. It is worth noting that our proposed method is also applicable to various paradigms of domain adaptive object detection. The code is available at:https://github.com/aiiu-lab/BlenDA
format Preprint
id arxiv_https___arxiv_org_abs_2401_09921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BlenDA: Domain Adaptive Object Detection through diffusion-based blending
Huang, Tzuhsuan
Huang, Chen-Che
Ku, Chung-Hao
Chen, Jun-Cheng
Computer Vision and Pattern Recognition
Unsupervised domain adaptation (UDA) aims to transfer a model learned using labeled data from the source domain to unlabeled data in the target domain. To address the large domain gap issue between the source and target domains, we propose a novel regularization method for domain adaptive object detection, BlenDA, by generating the pseudo samples of the intermediate domains and their corresponding soft domain labels for adaptation training. The intermediate samples are generated by dynamically blending the source images with their corresponding translated images using an off-the-shelf pre-trained text-to-image diffusion model which takes the text label of the target domain as input and has demonstrated superior image-to-image translation quality. Based on experimental results from two adaptation benchmarks, our proposed approach can significantly enhance the performance of the state-of-the-art domain adaptive object detector, Adversarial Query Transformer (AQT). Particularly, in the Cityscapes to Foggy Cityscapes adaptation, we achieve an impressive 53.4% mAP on the Foggy Cityscapes dataset, surpassing the previous state-of-the-art by 1.5%. It is worth noting that our proposed method is also applicable to various paradigms of domain adaptive object detection. The code is available at:https://github.com/aiiu-lab/BlenDA
title BlenDA: Domain Adaptive Object Detection through diffusion-based blending
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.09921