Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Qingcheng, Long, Pengyu, Zhang, Qixuan, Qin, Dafei, Liang, Han, Zhang, Longwen, Zhang, Yingliang, Yu, Jingyi, Xu, Lan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917577703292928
author Zhao, Qingcheng
Long, Pengyu
Zhang, Qixuan
Qin, Dafei
Liang, Han
Zhang, Longwen
Zhang, Yingliang
Yu, Jingyi
Xu, Lan
author_facet Zhao, Qingcheng
Long, Pengyu
Zhang, Qixuan
Qin, Dafei
Liang, Han
Zhang, Longwen
Zhang, Yingliang
Yu, Jingyi
Xu, Lan
contents The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism and a lack of lexible conditioning. We address this challenge through a trilogy. We first introduce Generalized Neural Parametric Facial Asset (GNPFA), an efficient variational auto-encoder mapping facial geometry and images to a highly generalized expression latent space, decoupling expressions and identities. Then, we utilize GNPFA to extract high-quality expressions and accurate head poses from a large array of videos. This presents the M2F-D dataset, a large, diverse, and scan-level co-speech 3D facial animation dataset with well-annotated emotional and style labels. Finally, we propose Media2Face, a diffusion model in GNPFA latent space for co-speech facial animation generation, accepting rich multi-modality guidances from audio, text, and image. Extensive experiments demonstrate that our model not only achieves high fidelity in facial animation synthesis but also broadens the scope of expressiveness and style adaptability in 3D facial animation.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15687
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
Zhao, Qingcheng
Long, Pengyu
Zhang, Qixuan
Qin, Dafei
Liang, Han
Zhang, Longwen
Zhang, Yingliang
Yu, Jingyi
Xu, Lan
Computer Vision and Pattern Recognition
Graphics
The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism and a lack of lexible conditioning. We address this challenge through a trilogy. We first introduce Generalized Neural Parametric Facial Asset (GNPFA), an efficient variational auto-encoder mapping facial geometry and images to a highly generalized expression latent space, decoupling expressions and identities. Then, we utilize GNPFA to extract high-quality expressions and accurate head poses from a large array of videos. This presents the M2F-D dataset, a large, diverse, and scan-level co-speech 3D facial animation dataset with well-annotated emotional and style labels. Finally, we propose Media2Face, a diffusion model in GNPFA latent space for co-speech facial animation generation, accepting rich multi-modality guidances from audio, text, and image. Extensive experiments demonstrate that our model not only achieves high fidelity in facial animation synthesis but also broadens the scope of expressiveness and style adaptability in 3D facial animation.
title Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2401.15687