One-for-All: Towards Universal Domain Translation with a Single StyleGAN

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Du, Yong, Zhan, Jiahui, Li, Xinzhe, Dong, Junyu, Chen, Sheng, Yang, Ming-Hsuan, He, Shengfeng
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917860194910208
author Du, Yong
Zhan, Jiahui
Li, Xinzhe
Dong, Junyu
Chen, Sheng
Yang, Ming-Hsuan
He, Shengfeng
author_facet Du, Yong
Zhan, Jiahui
Li, Xinzhe
Dong, Junyu
Chen, Sheng
Yang, Ming-Hsuan
He, Shengfeng
contents In this paper, we propose a novel translation model, UniTranslator, for transforming representations between visually distinct domains under conditions of limited training data and significant visual differences. The main idea behind our approach is leveraging the domain-neutral capabilities of CLIP as a bridging mechanism, while utilizing a separate module to extract abstract, domain-agnostic semantics from the embeddings of both the source and target realms. Fusing these abstract semantics with target-specific semantics results in a transformed embedding within the CLIP space. To bridge the gap between the disparate worlds of CLIP and StyleGAN, we introduce a new non-linear mapper, the CLIP2P mapper. Utilizing CLIP embeddings, this module is tailored to approximate the latent distribution in the StyleGAN's latent space, effectively acting as a connector between these two spaces. The proposed UniTranslator is versatile and capable of performing various tasks, including style mixing, stylization, and translations, even in visually challenging scenarios across different visual domains. Notably, UniTranslator generates high-quality translations that showcase domain relevance, diversity, and improved image quality. UniTranslator surpasses the performance of existing general-purpose models and performs well against specialized models in representative tasks. The source code and trained models will be released to the public.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14222
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle One-for-All: Towards Universal Domain Translation with a Single StyleGAN
Du, Yong
Zhan, Jiahui
Li, Xinzhe
Dong, Junyu
Chen, Sheng
Yang, Ming-Hsuan
He, Shengfeng
Computer Vision and Pattern Recognition
In this paper, we propose a novel translation model, UniTranslator, for transforming representations between visually distinct domains under conditions of limited training data and significant visual differences. The main idea behind our approach is leveraging the domain-neutral capabilities of CLIP as a bridging mechanism, while utilizing a separate module to extract abstract, domain-agnostic semantics from the embeddings of both the source and target realms. Fusing these abstract semantics with target-specific semantics results in a transformed embedding within the CLIP space. To bridge the gap between the disparate worlds of CLIP and StyleGAN, we introduce a new non-linear mapper, the CLIP2P mapper. Utilizing CLIP embeddings, this module is tailored to approximate the latent distribution in the StyleGAN's latent space, effectively acting as a connector between these two spaces. The proposed UniTranslator is versatile and capable of performing various tasks, including style mixing, stylization, and translations, even in visually challenging scenarios across different visual domains. Notably, UniTranslator generates high-quality translations that showcase domain relevance, diversity, and improved image quality. UniTranslator surpasses the performance of existing general-purpose models and performs well against specialized models in representative tasks. The source code and trained models will be released to the public.
title One-for-All: Towards Universal Domain Translation with a Single StyleGAN
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.14222