LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Xin, Zhuang, Chuanqing, Jin, Chenxi, Lu, Zhengda, Wang, Yiqun, Liu, Wu, Xiao, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914113934852096
author Lu, Xin
Zhuang, Chuanqing
Jin, Chenxi
Lu, Zhengda
Wang, Yiqun
Liu, Wu
Xiao, Jun
author_facet Lu, Xin
Zhuang, Chuanqing
Jin, Chenxi
Lu, Zhengda
Wang, Yiqun
Liu, Wu
Xiao, Jun
contents Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation, they still depend on explicit one-hot encodings to represent identity and emotion with given emotion and identity labels, which limits their ability to generalize to unseen speakers. Moreover, the emotional cues inherently present in speech are often neglected, limiting the naturalness and adaptability of generated animations. In this work, we propose LSF-Animation, a novel framework that eliminates the reliance on explicit emotion and identity feature representations. Specifically, LSF-Animation implicitly extracts emotion information from speech and captures the identity features from a neutral facial mesh, enabling improved generalization to unseen speakers and emotional states without requiring manual labels. Furthermore, we introduce a Hierarchical Interaction Fusion Block (HIFB), which employs a fusion token to integrate dual transformer features and more effectively integrate emotional, motion-related and identity-related cues. Extensive experiments conducted on the 3DMEAD dataset demonstrate that our method surpasses recent state-of-the-art approaches in terms of emotional expressiveness, identity generalization, and animation realism. The source code will be released at: https://github.com/Dogter521/LSF-Animation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation
Lu, Xin
Zhuang, Chuanqing
Jin, Chenxi
Lu, Zhengda
Wang, Yiqun
Liu, Wu
Xiao, Jun
Computer Vision and Pattern Recognition
Graphics
Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation, they still depend on explicit one-hot encodings to represent identity and emotion with given emotion and identity labels, which limits their ability to generalize to unseen speakers. Moreover, the emotional cues inherently present in speech are often neglected, limiting the naturalness and adaptability of generated animations. In this work, we propose LSF-Animation, a novel framework that eliminates the reliance on explicit emotion and identity feature representations. Specifically, LSF-Animation implicitly extracts emotion information from speech and captures the identity features from a neutral facial mesh, enabling improved generalization to unseen speakers and emotional states without requiring manual labels. Furthermore, we introduce a Hierarchical Interaction Fusion Block (HIFB), which employs a fusion token to integrate dual transformer features and more effectively integrate emotional, motion-related and identity-related cues. Extensive experiments conducted on the 3DMEAD dataset demonstrate that our method surpasses recent state-of-the-art approaches in terms of emotional expressiveness, identity generalization, and animation realism. The source code will be released at: https://github.com/Dogter521/LSF-Animation.
title LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2510.21864