PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Tianshun, Zhou, Benjia, Liu, Ajian, Liang, Yanyan, Zhang, Du, Lei, Zhen, Wan, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918232230723584
author Han, Tianshun
Zhou, Benjia
Liu, Ajian
Liang, Yanyan
Zhang, Du
Lei, Zhen
Wan, Jun
author_facet Han, Tianshun
Zhou, Benjia
Liu, Ajian
Liang, Yanyan
Zhang, Du
Lei, Zhen
Wan, Jun
contents PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introducing a Dual-Stream Emotion Extractor (DSEE) that captures both time and frequency-domain audio features for fine-grained emotion analysis, and an Emotional Style Modeling Module (ESMM) that models individual expression patterns based on voiceprint characteristics. To address data scarcity, the method leverages a newly constructed 3D-EmoStyle dataset. Evaluations demonstrate that PESTalk outperforms state-of-the-art methods in producing realistic and personalized facial animations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05121
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Han, Tianshun
Zhou, Benjia
Liu, Ajian
Liang, Yanyan
Zhang, Du
Lei, Zhen
Wan, Jun
Graphics
Artificial Intelligence
PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introducing a Dual-Stream Emotion Extractor (DSEE) that captures both time and frequency-domain audio features for fine-grained emotion analysis, and an Emotional Style Modeling Module (ESMM) that models individual expression patterns based on voiceprint characteristics. To address data scarcity, the method leverages a newly constructed 3D-EmoStyle dataset. Evaluations demonstrate that PESTalk outperforms state-of-the-art methods in producing realistic and personalized facial animations.
title PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
topic Graphics
Artificial Intelligence
url https://arxiv.org/abs/2512.05121