EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914193531207680 |
|---|---|
| author | Liu, Chang Jing, Tianjiao Ma, Chengcheng Zhou, Xuanqi Lian, Zhengxuan Jin, Qin Yuan, Hongliang Huang, Shi-Sheng |
| author_facet | Liu, Chang Jing, Tianjiao Ma, Chengcheng Zhou, Xuanqi Lian, Zhengxuan Jin, Qin Yuan, Hongliang Huang, Shi-Sheng |
| contents | Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This paper introduces a new editable 3D Gaussian talking head, i.e. EmoDiffTalk. Our key idea is a novel Emotion-aware Gaussian Diffusion, which includes an action unit (AU) prompt Gaussian diffusion process for fine-grained facial animator, and moreover an accurate text-to-AU emotion controller to provide accurate and expansive dynamic emotional editing using text input. Experiments on public EmoTalk3D and RenderMe-360 datasets demonstrate superior emotional subtlety, lip-sync fidelity, and controllability of our EmoDiffTalk over previous works, establishing a principled pathway toward high-quality, diffusion-driven, multimodal editable 3D talking-head synthesis. To our best knowledge, our EmoDiffTalk is one of the first few 3D Gaussian Splatting talking-head generation framework, especially supporting continuous, multimodal emotional editing within the AU-based expression space. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_05991 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head Liu, Chang Jing, Tianjiao Ma, Chengcheng Zhou, Xuanqi Lian, Zhengxuan Jin, Qin Yuan, Hongliang Huang, Shi-Sheng Computer Vision and Pattern Recognition Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This paper introduces a new editable 3D Gaussian talking head, i.e. EmoDiffTalk. Our key idea is a novel Emotion-aware Gaussian Diffusion, which includes an action unit (AU) prompt Gaussian diffusion process for fine-grained facial animator, and moreover an accurate text-to-AU emotion controller to provide accurate and expansive dynamic emotional editing using text input. Experiments on public EmoTalk3D and RenderMe-360 datasets demonstrate superior emotional subtlety, lip-sync fidelity, and controllability of our EmoDiffTalk over previous works, establishing a principled pathway toward high-quality, diffusion-driven, multimodal editable 3D talking-head synthesis. To our best knowledge, our EmoDiffTalk is one of the first few 3D Gaussian Splatting talking-head generation framework, especially supporting continuous, multimodal emotional editing within the AU-based expression space. |
| title | EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.05991 |