EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915222754689024 |
|---|---|
| author | Shen, Xuli Cai, Hua Yu, Dingding Shen, Weilin Xu, Qing Xue, Xiangyang |
| author_facet | Shen, Xuli Cai, Hua Yu, Dingding Shen, Weilin Xu, Qing Xue, Xiangyang |
| contents | Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_19416 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters Shen, Xuli Cai, Hua Yu, Dingding Shen, Weilin Xu, Qing Xue, Xiangyang Computer Vision and Pattern Recognition Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability. |
| title | EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.19416 |