EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shen, Xuli, Cai, Hua, Yu, Dingding, Shen, Weilin, Xu, Qing, Xue, Xiangyang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915222754689024
author Shen, Xuli
Cai, Hua
Yu, Dingding
Shen, Weilin
Xu, Qing
Xue, Xiangyang
author_facet Shen, Xuli
Cai, Hua
Yu, Dingding
Shen, Weilin
Xu, Qing
Xue, Xiangyang
contents Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19416
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters
Shen, Xuli
Cai, Hua
Yu, Dingding
Shen, Weilin
Xu, Qing
Xue, Xiangyang
Computer Vision and Pattern Recognition
Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability.
title EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.19416