DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Yuming, Xie, You, Xu, Hongyi, Song, Guoxian, Shi, Yichun, Chang, Di, Yang, Jing, Luo, Linjie
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917962305241088
author Gu, Yuming
Xie, You
Xu, Hongyi
Song, Guoxian
Shi, Yichun
Chang, Di
Yang, Jing
Luo, Linjie
author_facet Gu, Yuming
Xie, You
Xu, Hongyi
Song, Guoxian
Shi, Yichun
Chang, Di
Yang, Jing
Luo, Linjie
contents We present DiffPortrait3D, a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically, given a single RGB input, we aim to synthesize plausible but consistent facial details rendered from novel camera views with retained both identity and facial expression. In lieu of time-consuming optimization and fine-tuning, our zero-shot method generalizes well to arbitrary face portraits with unposed camera views, extreme facial expressions, and diverse artistic depictions. At its core, we leverage the generative prior of 2D diffusion models pre-trained on large-scale image datasets as our rendering backbone, while the denoising is guided with disentangled attentive control of appearance and camera pose. To achieve this, we first inject the appearance context from the reference image into the self-attention layers of the frozen UNets. The rendering view is then manipulated with a novel conditional control module that interprets the camera pose by watching a condition image of a crossed subject from the same view. Furthermore, we insert a trainable cross-view attention module to enhance view consistency, which is further strengthened with a novel 3D-aware noise generation process during inference. We demonstrate state-of-the-art results both qualitatively and quantitatively on our challenging in-the-wild and multi-view benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13016
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis
Gu, Yuming
Xie, You
Xu, Hongyi
Song, Guoxian
Shi, Yichun
Chang, Di
Yang, Jing
Luo, Linjie
Computer Vision and Pattern Recognition
We present DiffPortrait3D, a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically, given a single RGB input, we aim to synthesize plausible but consistent facial details rendered from novel camera views with retained both identity and facial expression. In lieu of time-consuming optimization and fine-tuning, our zero-shot method generalizes well to arbitrary face portraits with unposed camera views, extreme facial expressions, and diverse artistic depictions. At its core, we leverage the generative prior of 2D diffusion models pre-trained on large-scale image datasets as our rendering backbone, while the denoising is guided with disentangled attentive control of appearance and camera pose. To achieve this, we first inject the appearance context from the reference image into the self-attention layers of the frozen UNets. The rendering view is then manipulated with a novel conditional control module that interprets the camera pose by watching a condition image of a crossed subject from the same view. Furthermore, we insert a trainable cross-view attention module to enhance view consistency, which is further strengthened with a novel 3D-aware noise generation process during inference. We demonstrate state-of-the-art results both qualitatively and quantitatively on our challenging in-the-wild and multi-view benchmarks.
title DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.13016