Reference-Based 3D-Aware Image Editing with Triplanes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bilecen, Bahri Batuhan, Yalin, Yigit, Yu, Ning, Dundar, Aysegul
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913793694498816
author Bilecen, Bahri Batuhan
Yalin, Yigit
Yu, Ning
Dundar, Aysegul
author_facet Bilecen, Bahri Batuhan
Yalin, Yigit
Yu, Ning
Dundar, Aysegul
contents Generative Adversarial Networks (GANs) have emerged as powerful tools for high-quality image generation and real image editing by manipulating their latent spaces. Recent advancements in GANs include 3D-aware models such as EG3D, which feature efficient triplane-based architectures capable of reconstructing 3D geometry from single images. However, limited attention has been given to providing an integrated framework for 3D-aware, high-quality, reference-based image editing. This study addresses this gap by exploring and demonstrating the effectiveness of the triplane space for advanced reference-based edits. Our novel approach integrates encoding, automatic localization, spatial disentanglement of triplane features, and fusion learning to achieve the desired edits. We demonstrate how our approach excels across diverse domains, including human faces, 360-degree heads, animal faces, partially stylized edits like cartoon faces, full-body clothing edits, and edits on class-agnostic samples. Our method shows state-of-the-art performance over relevant latent direction, text, and image-guided 2D and 3D-aware diffusion and GAN methods, both qualitatively and quantitatively.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03632
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reference-Based 3D-Aware Image Editing with Triplanes
Bilecen, Bahri Batuhan
Yalin, Yigit
Yu, Ning
Dundar, Aysegul
Computer Vision and Pattern Recognition
Generative Adversarial Networks (GANs) have emerged as powerful tools for high-quality image generation and real image editing by manipulating their latent spaces. Recent advancements in GANs include 3D-aware models such as EG3D, which feature efficient triplane-based architectures capable of reconstructing 3D geometry from single images. However, limited attention has been given to providing an integrated framework for 3D-aware, high-quality, reference-based image editing. This study addresses this gap by exploring and demonstrating the effectiveness of the triplane space for advanced reference-based edits. Our novel approach integrates encoding, automatic localization, spatial disentanglement of triplane features, and fusion learning to achieve the desired edits. We demonstrate how our approach excels across diverse domains, including human faces, 360-degree heads, animal faces, partially stylized edits like cartoon faces, full-body clothing edits, and edits on class-agnostic samples. Our method shows state-of-the-art performance over relevant latent direction, text, and image-guided 2D and 3D-aware diffusion and GAN methods, both qualitatively and quantitatively.
title Reference-Based 3D-Aware Image Editing with Triplanes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.03632