EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Chang, Jing, Tianjiao, Ma, Chengcheng, Zhou, Xuanqi, Lian, Zhengxuan, Jin, Qin, Yuan, Hongliang, Huang, Shi-Sheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914193531207680
author Liu, Chang
Jing, Tianjiao
Ma, Chengcheng
Zhou, Xuanqi
Lian, Zhengxuan
Jin, Qin
Yuan, Hongliang
Huang, Shi-Sheng
author_facet Liu, Chang
Jing, Tianjiao
Ma, Chengcheng
Zhou, Xuanqi
Lian, Zhengxuan
Jin, Qin
Yuan, Hongliang
Huang, Shi-Sheng
contents Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This paper introduces a new editable 3D Gaussian talking head, i.e. EmoDiffTalk. Our key idea is a novel Emotion-aware Gaussian Diffusion, which includes an action unit (AU) prompt Gaussian diffusion process for fine-grained facial animator, and moreover an accurate text-to-AU emotion controller to provide accurate and expansive dynamic emotional editing using text input. Experiments on public EmoTalk3D and RenderMe-360 datasets demonstrate superior emotional subtlety, lip-sync fidelity, and controllability of our EmoDiffTalk over previous works, establishing a principled pathway toward high-quality, diffusion-driven, multimodal editable 3D talking-head synthesis. To our best knowledge, our EmoDiffTalk is one of the first few 3D Gaussian Splatting talking-head generation framework, especially supporting continuous, multimodal emotional editing within the AU-based expression space.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
Liu, Chang
Jing, Tianjiao
Ma, Chengcheng
Zhou, Xuanqi
Lian, Zhengxuan
Jin, Qin
Yuan, Hongliang
Huang, Shi-Sheng
Computer Vision and Pattern Recognition
Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This paper introduces a new editable 3D Gaussian talking head, i.e. EmoDiffTalk. Our key idea is a novel Emotion-aware Gaussian Diffusion, which includes an action unit (AU) prompt Gaussian diffusion process for fine-grained facial animator, and moreover an accurate text-to-AU emotion controller to provide accurate and expansive dynamic emotional editing using text input. Experiments on public EmoTalk3D and RenderMe-360 datasets demonstrate superior emotional subtlety, lip-sync fidelity, and controllability of our EmoDiffTalk over previous works, establishing a principled pathway toward high-quality, diffusion-driven, multimodal editable 3D talking-head synthesis. To our best knowledge, our EmoDiffTalk is one of the first few 3D Gaussian Splatting talking-head generation framework, especially supporting continuous, multimodal emotional editing within the AU-based expression space.
title EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05991