Learning to Generate Conditional Tri-plane for 3D-aware Expression Controllable Portrait Animation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ki, Taekyung, Min, Dongchan, Chae, Gyeongsu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917310729551872
author Ki, Taekyung
Min, Dongchan
Chae, Gyeongsu
author_facet Ki, Taekyung
Min, Dongchan
Chae, Gyeongsu
contents In this paper, we present Export3D, a one-shot 3D-aware portrait animation method that is able to control the facial expression and camera view of a given portrait image. To achieve this, we introduce a tri-plane generator with an effective expression conditioning method, which directly generates a tri-plane of 3D prior by transferring the expression parameter of 3DMM into the source image. The tri-plane is then decoded into the image of different view through a differentiable volume rendering. Existing portrait animation methods heavily rely on image warping to transfer the expression in the motion space, challenging on disentanglement of appearance and expression. In contrast, we propose a contrastive pre-training framework for appearance-free expression parameter, eliminating undesirable appearance swap when transferring a cross-identity expression. Extensive experiments show that our pre-training framework can learn the appearance-free expression representation hidden in 3DMM, and our model can generate 3D-aware expression controllable portrait images without appearance swap in the cross-identity manner.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00636
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Generate Conditional Tri-plane for 3D-aware Expression Controllable Portrait Animation
Ki, Taekyung
Min, Dongchan
Chae, Gyeongsu
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
In this paper, we present Export3D, a one-shot 3D-aware portrait animation method that is able to control the facial expression and camera view of a given portrait image. To achieve this, we introduce a tri-plane generator with an effective expression conditioning method, which directly generates a tri-plane of 3D prior by transferring the expression parameter of 3DMM into the source image. The tri-plane is then decoded into the image of different view through a differentiable volume rendering. Existing portrait animation methods heavily rely on image warping to transfer the expression in the motion space, challenging on disentanglement of appearance and expression. In contrast, we propose a contrastive pre-training framework for appearance-free expression parameter, eliminating undesirable appearance swap when transferring a cross-identity expression. Extensive experiments show that our pre-training framework can learn the appearance-free expression representation hidden in 3DMM, and our model can generate 3D-aware expression controllable portrait images without appearance swap in the cross-identity manner.
title Learning to Generate Conditional Tri-plane for 3D-aware Expression Controllable Portrait Animation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2404.00636