EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Qianyun, Ji, Xinya, Gong, Yicheng, Lu, Yuanxun, Diao, Zhengyu, Huang, Linjia, Yao, Yao, Zhu, Siyu, Ma, Zhan, Xu, Songcen, Wu, Xiaofei, Zhang, Zixiao, Cao, Xun, Zhu, Hao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914895510896640
author He, Qianyun
Ji, Xinya
Gong, Yicheng
Lu, Yuanxun
Diao, Zhengyu
Huang, Linjia
Yao, Yao
Zhu, Siyu
Ma, Zhan
Xu, Songcen
Wu, Xiaofei
Zhang, Zixiao
Cao, Xun
Zhu, Hao
author_facet He, Qianyun
Ji, Xinya
Gong, Yicheng
Lu, Yuanxun
Diao, Zhengyu
Huang, Linjia
Yao, Yao
Zhu, Siyu
Ma, Zhan
Xu, Songcen
Wu, Xiaofei
Zhang, Zixiao
Cao, Xun
Zhu, Hao
contents We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view consistency and a lack of emotional expressiveness. To address these issues, we collect EmoTalk3D dataset with calibrated multi-view videos, emotional annotations, and per-frame 3D geometry. By training on the EmoTalk3D dataset, we propose a \textit{`Speech-to-Geometry-to-Appearance'} mapping framework that first predicts faithful 3D geometry sequence from the audio features, then the appearance of a 3D talking head represented by 4D Gaussians is synthesized from the predicted geometry. The appearance is further disentangled into canonical and dynamic Gaussians, learned from multi-view videos, and fused to render free-view talking head animation. Moreover, our model enables controllable emotion in the generated talking heads and can be rendered in wide-range views. Our method exhibits improved rendering quality and stability in lip motion generation while capturing dynamic facial details such as wrinkles and subtle expressions. Experiments demonstrate the effectiveness of our approach in generating high-fidelity and emotion-controllable 3D talking heads. The code and EmoTalk3D dataset are released at https://nju-3dv.github.io/projects/EmoTalk3D.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00297
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head
He, Qianyun
Ji, Xinya
Gong, Yicheng
Lu, Yuanxun
Diao, Zhengyu
Huang, Linjia
Yao, Yao
Zhu, Siyu
Ma, Zhan
Xu, Songcen
Wu, Xiaofei
Zhang, Zixiao
Cao, Xun
Zhu, Hao
Computer Vision and Pattern Recognition
We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view consistency and a lack of emotional expressiveness. To address these issues, we collect EmoTalk3D dataset with calibrated multi-view videos, emotional annotations, and per-frame 3D geometry. By training on the EmoTalk3D dataset, we propose a \textit{`Speech-to-Geometry-to-Appearance'} mapping framework that first predicts faithful 3D geometry sequence from the audio features, then the appearance of a 3D talking head represented by 4D Gaussians is synthesized from the predicted geometry. The appearance is further disentangled into canonical and dynamic Gaussians, learned from multi-view videos, and fused to render free-view talking head animation. Moreover, our model enables controllable emotion in the generated talking heads and can be rendered in wide-range views. Our method exhibits improved rendering quality and stability in lip motion generation while capturing dynamic facial details such as wrinkles and subtle expressions. Experiments demonstrate the effectiveness of our approach in generating high-fidelity and emotion-controllable 3D talking heads. The code and EmoTalk3D dataset are released at https://nju-3dv.github.io/projects/EmoTalk3D.
title EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.00297