UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hengjia, Liu, Yang, Lin, Yuqi, Zhang, Zhanwei, Zhao, Yibo, Pan, weihang, Zheng, Tu, Yang, Zheng, Jiang, Yuchun, Wu, Boxi, Cai, Deng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911798312042496
author Li, Hengjia
Liu, Yang
Lin, Yuqi
Zhang, Zhanwei
Zhao, Yibo
Pan, weihang
Zheng, Tu
Yang, Zheng
Jiang, Yuchun
Wu, Boxi
Cai, Deng
author_facet Li, Hengjia
Liu, Yang
Lin, Yuqi
Zhang, Zhanwei
Zhao, Yibo
Pan, weihang
Zheng, Tu
Yang, Zheng
Jiang, Yuchun
Wu, Boxi
Cai, Deng
contents Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt the generator to a single target domain and are limited to a single modality, either text-driven or image-driven. Moreover, they cannot maintain well consistency with the source domain, which impedes the inheritance of the diversity. In this paper, we propose UniHDA, a \textbf{unified} and \textbf{versatile} framework for generative hybrid domain adaptation with multi-modal references from multiple domains. We use CLIP encoder to project multi-modal references into a unified embedding space and then linearly interpolate the direction vectors from multiple target domains to achieve hybrid domain adaptation. To ensure \textbf{consistency} with the source domain, we propose a novel cross-domain spatial structure (CSS) loss that maintains detailed spatial structure information between source and target generator. Experiments show that the adapted generator can synthesise realistic images with various attribute compositions. Additionally, our framework is generator-agnostic and versatile to multiple generators, e.g., StyleGAN, EG3D, and Diffusion Models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation
Li, Hengjia
Liu, Yang
Lin, Yuqi
Zhang, Zhanwei
Zhao, Yibo
Pan, weihang
Zheng, Tu
Yang, Zheng
Jiang, Yuchun
Wu, Boxi
Cai, Deng
Computer Vision and Pattern Recognition
Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt the generator to a single target domain and are limited to a single modality, either text-driven or image-driven. Moreover, they cannot maintain well consistency with the source domain, which impedes the inheritance of the diversity. In this paper, we propose UniHDA, a \textbf{unified} and \textbf{versatile} framework for generative hybrid domain adaptation with multi-modal references from multiple domains. We use CLIP encoder to project multi-modal references into a unified embedding space and then linearly interpolate the direction vectors from multiple target domains to achieve hybrid domain adaptation. To ensure \textbf{consistency} with the source domain, we propose a novel cross-domain spatial structure (CSS) loss that maintains detailed spatial structure information between source and target generator. Experiments show that the adapted generator can synthesise realistic images with various attribute compositions. Additionally, our framework is generator-agnostic and versatile to multiple generators, e.g., StyleGAN, EG3D, and Diffusion Models.
title UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.12596