KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Zhihao, Gong, Shengjie, Tang, Jiapeng, Liang, Lingyu, Huang, Yining, Li, Haojie, Huang, Shuangping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916377479086080
author Xu, Zhihao
Gong, Shengjie
Tang, Jiapeng
Liang, Lingyu
Huang, Yining
Li, Haojie
Huang, Shuangping
author_facet Xu, Zhihao
Gong, Shengjie
Tang, Jiapeng
Liang, Lingyu
Huang, Yining
Li, Haojie
Huang, Shuangping
contents We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains challenging. Direct regression of the entire sequence often leads to over-smoothed results due to the ill-posed nature of the problem. To this end, we propose a progressive learning mechanism that generates 3D facial animations by introducing key motion capture to decrease cross-modal mapping uncertainty and learning complexity. Concretely, our method integrates linguistic and data-driven priors through two modules: the linguistic-based key motion acquisition and the cross-modal motion completion. The former identifies key motions and learns the associated 3D facial expressions, ensuring accurate lip-speech synchronization. The latter extends key motions into a full sequence of 3D talking faces guided by audio features, improving temporal coherence and audio-visual consistency. Extensive experimental comparisons against existing state-of-the-art methods demonstrate the superiority of our approach in generating more vivid and consistent talking face animations. Consistent enhancements in results through the integration of our proposed learning scheme with existing methods underscore the efficacy of our approach. Our code and weights will be at the project website: \url{https://github.com/ffxzh/KMTalk}.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding
Xu, Zhihao
Gong, Shengjie
Tang, Jiapeng
Liang, Lingyu
Huang, Yining
Li, Haojie
Huang, Shuangping
Computer Vision and Pattern Recognition
We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains challenging. Direct regression of the entire sequence often leads to over-smoothed results due to the ill-posed nature of the problem. To this end, we propose a progressive learning mechanism that generates 3D facial animations by introducing key motion capture to decrease cross-modal mapping uncertainty and learning complexity. Concretely, our method integrates linguistic and data-driven priors through two modules: the linguistic-based key motion acquisition and the cross-modal motion completion. The former identifies key motions and learns the associated 3D facial expressions, ensuring accurate lip-speech synchronization. The latter extends key motions into a full sequence of 3D talking faces guided by audio features, improving temporal coherence and audio-visual consistency. Extensive experimental comparisons against existing state-of-the-art methods demonstrate the superiority of our approach in generating more vivid and consistent talking face animations. Consistent enhancements in results through the integration of our proposed learning scheme with existing methods underscore the efficacy of our approach. Our code and weights will be at the project website: \url{https://github.com/ffxzh/KMTalk}.
title KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.01113