EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Bingyuan, Zhang, Xulong, Cheng, Ning, Yu, Jun, Xiao, Jing, Wang, Jianzong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911758516486144
author Zhang, Bingyuan
Zhang, Xulong
Cheng, Ning
Yu, Jun
Xiao, Jing
Wang, Jianzong
author_facet Zhang, Bingyuan
Zhang, Xulong
Cheng, Ning
Yu, Jun
Xiao, Jing
Wang, Jianzong
contents In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with challenging identities. Furthermore, methods for editing expressions are often confined to a singular emotion, failing to adapt to intricate emotions. To overcome these challenges, this paper proposes EmoTalker, an emotionally editable portraits animation approach based on the diffusion model. EmoTalker modifies the denoising process to ensure preservation of the original portrait's identity during inference. To enhance emotion comprehension from text input, Emotion Intensity Block is introduced to analyze fine-grained emotions and strengths derived from prompts. Additionally, a crafted dataset is harnessed to enhance emotion comprehension within prompts. Experiments show the effectiveness of EmoTalker in generating high-quality, emotionally customizable facial expressions.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
Zhang, Bingyuan
Zhang, Xulong
Cheng, Ning
Yu, Jun
Xiao, Jing
Wang, Jianzong
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with challenging identities. Furthermore, methods for editing expressions are often confined to a singular emotion, failing to adapt to intricate emotions. To overcome these challenges, this paper proposes EmoTalker, an emotionally editable portraits animation approach based on the diffusion model. EmoTalker modifies the denoising process to ensure preservation of the original portrait's identity during inference. To enhance emotion comprehension from text input, Emotion Intensity Block is introduced to analyze fine-grained emotions and strengths derived from prompts. Additionally, a crafted dataset is harnessed to enhance emotion comprehension within prompts. Experiments show the effectiveness of EmoTalker in generating high-quality, emotionally customizable facial expressions.
title EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
topic Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.08049