HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anees, Abdul Basit, Baykal, Ahmet Canberk, Kizil, Muhammed Burak, Ceylan, Duygu, Erdem, Erkut, Erdem, Aykut
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913581444890624
author Anees, Abdul Basit
Baykal, Ahmet Canberk
Kizil, Muhammed Burak
Ceylan, Duygu
Erdem, Erkut
Erdem, Aykut
author_facet Anees, Abdul Basit
Baykal, Ahmet Canberk
Kizil, Muhammed Burak
Ceylan, Duygu
Erdem, Erkut
Erdem, Aykut
contents Generative Adversarial Networks (GANs), particularly StyleGAN and its variants, have demonstrated remarkable capabilities in generating highly realistic images. Despite their success, adapting these models to diverse tasks such as domain adaptation, reference-guided synthesis, and text-guided manipulation with limited training data remains challenging. Towards this end, in this study, we present a novel framework that significantly extends the capabilities of a pre-trained StyleGAN by integrating CLIP space via hypernetworks. This integration allows dynamic adaptation of StyleGAN to new domains defined by reference images or textual descriptions. Additionally, we introduce a CLIP-guided discriminator that enhances the alignment between generated images and target domains, ensuring superior image quality. Our approach demonstrates unprecedented flexibility, enabling text-guided image manipulation without the need for text-specific training data and facilitating seamless style transfer. Comprehensive qualitative and quantitative evaluations confirm the robustness and superior performance of our framework compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12832
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
Anees, Abdul Basit
Baykal, Ahmet Canberk
Kizil, Muhammed Burak
Ceylan, Duygu
Erdem, Erkut
Erdem, Aykut
Computer Vision and Pattern Recognition
Generative Adversarial Networks (GANs), particularly StyleGAN and its variants, have demonstrated remarkable capabilities in generating highly realistic images. Despite their success, adapting these models to diverse tasks such as domain adaptation, reference-guided synthesis, and text-guided manipulation with limited training data remains challenging. Towards this end, in this study, we present a novel framework that significantly extends the capabilities of a pre-trained StyleGAN by integrating CLIP space via hypernetworks. This integration allows dynamic adaptation of StyleGAN to new domains defined by reference images or textual descriptions. Additionally, we introduce a CLIP-guided discriminator that enhances the alignment between generated images and target domains, ensuring superior image quality. Our approach demonstrates unprecedented flexibility, enabling text-guided image manipulation without the need for text-specific training data and facilitating seamless style transfer. Comprehensive qualitative and quantitative evaluations confirm the robustness and superior performance of our framework compared to existing methods.
title HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.12832