PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918232230723584 |
|---|---|
| author | Han, Tianshun Zhou, Benjia Liu, Ajian Liang, Yanyan Zhang, Du Lei, Zhen Wan, Jun |
| author_facet | Han, Tianshun Zhou, Benjia Liu, Ajian Liang, Yanyan Zhang, Du Lei, Zhen Wan, Jun |
| contents | PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introducing a Dual-Stream Emotion Extractor (DSEE) that captures both time and frequency-domain audio features for fine-grained emotion analysis, and an Emotional Style Modeling Module (ESMM) that models individual expression patterns based on voiceprint characteristics. To address data scarcity, the method leverages a newly constructed 3D-EmoStyle dataset. Evaluations demonstrate that PESTalk outperforms state-of-the-art methods in producing realistic and personalized facial animations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_05121 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles Han, Tianshun Zhou, Benjia Liu, Ajian Liang, Yanyan Zhang, Du Lei, Zhen Wan, Jun Graphics Artificial Intelligence PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introducing a Dual-Stream Emotion Extractor (DSEE) that captures both time and frequency-domain audio features for fine-grained emotion analysis, and an Emotional Style Modeling Module (ESMM) that models individual expression patterns based on voiceprint characteristics. To address data scarcity, the method leverages a newly constructed 3D-EmoStyle dataset. Evaluations demonstrate that PESTalk outperforms state-of-the-art methods in producing realistic and personalized facial animations. |
| title | PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles |
| topic | Graphics Artificial Intelligence |
| url | https://arxiv.org/abs/2512.05121 |