3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sha, Xuanmeng, Zhang, Liyun, Mashita, Tomohiro, Chiba, Naoya, Uranishi, Yuki
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918123283677184
author Sha, Xuanmeng
Zhang, Liyun
Mashita, Tomohiro
Chiba, Naoya
Uranishi, Yuki
author_facet Sha, Xuanmeng
Zhang, Liyun
Mashita, Tomohiro
Chiba, Naoya
Uranishi, Yuki
contents Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10848
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
Sha, Xuanmeng
Zhang, Liyun
Mashita, Tomohiro
Chiba, Naoya
Uranishi, Yuki
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Sound
Audio and Speech Processing
Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations.
title 3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.10848