Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Guanwen, Cheng, Haoran, Li, Yunan, Ma, Zhiyuan, Li, Chaoneng, Qian, Zhihao, Miao, Qiguang, Pun, Chi-Man
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2402.01422
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913221320900608
author Feng, Guanwen
Cheng, Haoran
Li, Yunan
Ma, Zhiyuan
Li, Chaoneng
Qian, Zhihao
Miao, Qiguang
Pun, Chi-Man
author_facet Feng, Guanwen
Cheng, Haoran
Li, Yunan
Ma, Zhiyuan
Li, Chaoneng
Qian, Zhihao
Miao, Qiguang
Pun, Chi-Man
contents Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page: https://peterfanfan.github.io/EmoSpeaker/
format Preprint
id arxiv_https___arxiv_org_abs_2402_01422
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face Generation
Feng, Guanwen
Cheng, Haoran
Li, Yunan
Ma, Zhiyuan
Li, Chaoneng
Qian, Zhihao
Miao, Qiguang
Pun, Chi-Man
Computer Vision and Pattern Recognition
Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page: https://peterfanfan.github.io/EmoSpeaker/
title EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.01422