InstructPix2NeRF: Instructed 3D Portrait Editing from a Single Image

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jianhui, Liu, Shilong, Liu, Zidong, Wang, Yikai, Zheng, Kaiwen, Xu, Jinghui, Li, Jianmin, Zhu, Jun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910315579441152
author Li, Jianhui
Liu, Shilong
Liu, Zidong
Wang, Yikai
Zheng, Kaiwen
Xu, Jinghui
Li, Jianmin
Zhu, Jun
author_facet Li, Jianhui
Liu, Shilong
Liu, Zidong
Wang, Yikai
Zheng, Kaiwen
Xu, Jinghui
Li, Jianmin
Zhu, Jun
contents With the success of Neural Radiance Field (NeRF) in 3D-aware portrait editing, a variety of works have achieved promising results regarding both quality and 3D consistency. However, these methods heavily rely on per-prompt optimization when handling natural language as editing instructions. Due to the lack of labeled human face 3D datasets and effective architectures, the area of human-instructed 3D-aware editing for open-world portraits in an end-to-end manner remains under-explored. To solve this problem, we propose an end-to-end diffusion-based framework termed InstructPix2NeRF, which enables instructed 3D-aware portrait editing from a single open-world image with human instructions. At its core lies a conditional latent 3D diffusion process that lifts 2D editing to 3D space by learning the correlation between the paired images' difference and the instructions via triplet data. With the help of our proposed token position randomization strategy, we could even achieve multi-semantic editing through one single pass with the portrait identity well-preserved. Besides, we further propose an identity consistency module that directly modulates the extracted identity signals into our diffusion process, which increases the multi-view 3D identity consistency. Extensive experiments verify the effectiveness of our method and show its superiority against strong baselines quantitatively and qualitatively. Source code and pre-trained models can be found on our project page: \url{https://mybabyyh.github.io/InstructPix2NeRF}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_02826
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle InstructPix2NeRF: Instructed 3D Portrait Editing from a Single Image
Li, Jianhui
Liu, Shilong
Liu, Zidong
Wang, Yikai
Zheng, Kaiwen
Xu, Jinghui
Li, Jianmin
Zhu, Jun
Computer Vision and Pattern Recognition
With the success of Neural Radiance Field (NeRF) in 3D-aware portrait editing, a variety of works have achieved promising results regarding both quality and 3D consistency. However, these methods heavily rely on per-prompt optimization when handling natural language as editing instructions. Due to the lack of labeled human face 3D datasets and effective architectures, the area of human-instructed 3D-aware editing for open-world portraits in an end-to-end manner remains under-explored. To solve this problem, we propose an end-to-end diffusion-based framework termed InstructPix2NeRF, which enables instructed 3D-aware portrait editing from a single open-world image with human instructions. At its core lies a conditional latent 3D diffusion process that lifts 2D editing to 3D space by learning the correlation between the paired images' difference and the instructions via triplet data. With the help of our proposed token position randomization strategy, we could even achieve multi-semantic editing through one single pass with the portrait identity well-preserved. Besides, we further propose an identity consistency module that directly modulates the extracted identity signals into our diffusion process, which increases the multi-view 3D identity consistency. Extensive experiments verify the effectiveness of our method and show its superiority against strong baselines quantitatively and qualitatively. Source code and pre-trained models can be found on our project page: \url{https://mybabyyh.github.io/InstructPix2NeRF}.
title InstructPix2NeRF: Instructed 3D Portrait Editing from a Single Image
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.02826