Coherent3D: Coherent 3D Portrait Video Reconstruction via Triplane Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shengze, Li, Xueting, Liu, Chao, Chan, Matthew, Stengel, Michael, Fuchs, Henry, De Mello, Shalini, Nagano, Koki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915060566196224
author Wang, Shengze
Li, Xueting
Liu, Chao
Chan, Matthew
Stengel, Michael
Fuchs, Henry
De Mello, Shalini
Nagano, Koki
author_facet Wang, Shengze
Li, Xueting
Liu, Chao
Chan, Matthew
Stengel, Michael
Fuchs, Henry
De Mello, Shalini
Nagano, Koki
contents Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user's appearance. On the other hand, self-reenactment methods can render coherent 3D portraits by driving a 3D avatar built from a single reference image, but fail to faithfully preserve the user's per-frame appearance (e.g., instantaneous facial expression and lighting). As a result, none of these two frameworks is an ideal solution for democratized 3D telepresence. In this work, we address this dilemma and propose a novel solution that maintains both coherent identity and dynamic per-frame appearance to enable the best possible realism. To this end, we propose a new fusion-based method that takes the best of both worlds by fusing a canonical 3D prior from a reference view with dynamic appearance from per-frame input views, producing temporally stable 3D videos with faithful reconstruction of the user's per-frame appearance. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction and temporal consistency on in-studio and in-the-wild datasets. https://research.nvidia.com/labs/amri/projects/coherent3d
format Preprint
id arxiv_https___arxiv_org_abs_2412_08684
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Coherent3D: Coherent 3D Portrait Video Reconstruction via Triplane Fusion
Wang, Shengze
Li, Xueting
Liu, Chao
Chan, Matthew
Stengel, Michael
Fuchs, Henry
De Mello, Shalini
Nagano, Koki
Computer Vision and Pattern Recognition
Image and Video Processing
Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user's appearance. On the other hand, self-reenactment methods can render coherent 3D portraits by driving a 3D avatar built from a single reference image, but fail to faithfully preserve the user's per-frame appearance (e.g., instantaneous facial expression and lighting). As a result, none of these two frameworks is an ideal solution for democratized 3D telepresence. In this work, we address this dilemma and propose a novel solution that maintains both coherent identity and dynamic per-frame appearance to enable the best possible realism. To this end, we propose a new fusion-based method that takes the best of both worlds by fusing a canonical 3D prior from a reference view with dynamic appearance from per-frame input views, producing temporally stable 3D videos with faithful reconstruction of the user's per-frame appearance. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction and temporal consistency on in-studio and in-the-wild datasets. https://research.nvidia.com/labs/amri/projects/coherent3d
title Coherent3D: Coherent 3D Portrait Video Reconstruction via Triplane Fusion
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2412.08684