MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ye, Zhenhui, Zhong, Tianyun, Ren, Yi, Jiang, Ziyue, Huang, Jiawei, Huang, Rongjie, Liu, Jinglin, He, Jinzheng, Zhang, Chen, Wang, Zehan, Chen, Xize, Yin, Xiang, Zhao, Zhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912073114451968
author Ye, Zhenhui
Zhong, Tianyun
Ren, Yi
Jiang, Ziyue
Huang, Jiawei
Huang, Rongjie
Liu, Jinglin
He, Jinzheng
Zhang, Chen
Wang, Zehan
Chen, Xize
Yin, Xiang
Zhao, Zhou
author_facet Ye, Zhenhui
Zhong, Tianyun
Ren, Yi
Jiang, Ziyue
Huang, Jiawei
Huang, Rongjie
Liu, Jinglin
He, Jinzheng
Zhang, Chen
Wang, Zehan
Chen, Xize
Yin, Xiang
Zhao, Zhou
contents Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Source code and video samples are available at https://mimictalk.github.io .
format Preprint
id arxiv_https___arxiv_org_abs_2410_06734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
Ye, Zhenhui
Zhong, Tianyun
Ren, Yi
Jiang, Ziyue
Huang, Jiawei
Huang, Rongjie
Liu, Jinglin
He, Jinzheng
Zhang, Chen
Wang, Zehan
Chen, Xize
Yin, Xiang
Zhao, Zhou
Computer Vision and Pattern Recognition
Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Source code and video samples are available at https://mimictalk.github.io .
title MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.06734