UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911798312042496 |
|---|---|
| author | Li, Hengjia Liu, Yang Lin, Yuqi Zhang, Zhanwei Zhao, Yibo Pan, weihang Zheng, Tu Yang, Zheng Jiang, Yuchun Wu, Boxi Cai, Deng |
| author_facet | Li, Hengjia Liu, Yang Lin, Yuqi Zhang, Zhanwei Zhao, Yibo Pan, weihang Zheng, Tu Yang, Zheng Jiang, Yuchun Wu, Boxi Cai, Deng |
| contents | Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt the generator to a single target domain and are limited to a single modality, either text-driven or image-driven. Moreover, they cannot maintain well consistency with the source domain, which impedes the inheritance of the diversity. In this paper, we propose UniHDA, a \textbf{unified} and \textbf{versatile} framework for generative hybrid domain adaptation with multi-modal references from multiple domains. We use CLIP encoder to project multi-modal references into a unified embedding space and then linearly interpolate the direction vectors from multiple target domains to achieve hybrid domain adaptation. To ensure \textbf{consistency} with the source domain, we propose a novel cross-domain spatial structure (CSS) loss that maintains detailed spatial structure information between source and target generator. Experiments show that the adapted generator can synthesise realistic images with various attribute compositions. Additionally, our framework is generator-agnostic and versatile to multiple generators, e.g., StyleGAN, EG3D, and Diffusion Models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_12596 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation Li, Hengjia Liu, Yang Lin, Yuqi Zhang, Zhanwei Zhao, Yibo Pan, weihang Zheng, Tu Yang, Zheng Jiang, Yuchun Wu, Boxi Cai, Deng Computer Vision and Pattern Recognition Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt the generator to a single target domain and are limited to a single modality, either text-driven or image-driven. Moreover, they cannot maintain well consistency with the source domain, which impedes the inheritance of the diversity. In this paper, we propose UniHDA, a \textbf{unified} and \textbf{versatile} framework for generative hybrid domain adaptation with multi-modal references from multiple domains. We use CLIP encoder to project multi-modal references into a unified embedding space and then linearly interpolate the direction vectors from multiple target domains to achieve hybrid domain adaptation. To ensure \textbf{consistency} with the source domain, we propose a novel cross-domain spatial structure (CSS) loss that maintains detailed spatial structure information between source and target generator. Experiments show that the adapted generator can synthesise realistic images with various attribute compositions. Additionally, our framework is generator-agnostic and versatile to multiple generators, e.g., StyleGAN, EG3D, and Diffusion Models. |
| title | UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2401.12596 |