3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918123283677184 |
|---|---|
| author | Sha, Xuanmeng Zhang, Liyun Mashita, Tomohiro Chiba, Naoya Uranishi, Yuki |
| author_facet | Sha, Xuanmeng Zhang, Liyun Mashita, Tomohiro Chiba, Naoya Uranishi, Yuki |
| contents | Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_10848 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | 3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control Sha, Xuanmeng Zhang, Liyun Mashita, Tomohiro Chiba, Naoya Uranishi, Yuki Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Multimedia Sound Audio and Speech Processing Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations. |
| title | 3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Multimedia Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.10848 |