EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haotian, Weng, Yuzhe, Li, Yueyan, Guo, Zilu, Du, Jun, Niu, Shutong, Ma, Jiefeng, He, Shan, Wu, Xiaoyan, Hu, Qiming, Yin, Bing, Liu, Cong, Liu, Qingfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917869773651968
author Wang, Haotian
Weng, Yuzhe
Li, Yueyan
Guo, Zilu
Du, Jun
Niu, Shutong
Ma, Jiefeng
He, Shan
Wu, Xiaoyan
Hu, Qiming
Yin, Bing
Liu, Cong
Liu, Qingfeng
author_facet Wang, Haotian
Weng, Yuzhe
Li, Yueyan
Guo, Zilu
Du, Jun
Niu, Shutong
Ma, Jiefeng
He, Shan
Wu, Xiaoyan
Hu, Qiming
Yin, Bing
Liu, Cong
Liu, Qingfeng
contents Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address these issues. Firstly, to realize better control over the generation of lip movement and facial expression, a Vision-guided Audio Information Decoupling (V-AID) approach is designed to generate audio-based decoupled representations aligned with lip movements and expression. Specifically, to achieve alignment between audio and facial expression representation spaces, we present a Diffusion-based Co-speech Temporal Expansion (Di-CTE) module within V-AID to generate expression-related representations under multi-source emotion condition constraints. Then we propose a well-designed Emotional Talking Head Diffusion (ETHD) backbone to efficiently generate highly expressive talking head videos, which contains an Expression Decoupling Injection (EDI) module to automatically decouple the expressions from reference portraits while integrating the target expression information, achieving more expressive generation performance. Experimental results show that EmotiveTalk can generate expressive talking head videos, ensuring the promised controllability of emotions and stability during long-time generation, yielding state-of-the-art performance compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16726
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
Wang, Haotian
Weng, Yuzhe
Li, Yueyan
Guo, Zilu
Du, Jun
Niu, Shutong
Ma, Jiefeng
He, Shan
Wu, Xiaoyan
Hu, Qiming
Yin, Bing
Liu, Cong
Liu, Qingfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address these issues. Firstly, to realize better control over the generation of lip movement and facial expression, a Vision-guided Audio Information Decoupling (V-AID) approach is designed to generate audio-based decoupled representations aligned with lip movements and expression. Specifically, to achieve alignment between audio and facial expression representation spaces, we present a Diffusion-based Co-speech Temporal Expansion (Di-CTE) module within V-AID to generate expression-related representations under multi-source emotion condition constraints. Then we propose a well-designed Emotional Talking Head Diffusion (ETHD) backbone to efficiently generate highly expressive talking head videos, which contains an Expression Decoupling Injection (EDI) module to automatically decouple the expressions from reference portraits while integrating the target expression information, achieving more expressive generation performance. Experimental results show that EmotiveTalk can generate expressive talking head videos, ensuring the promised controllability of emotions and stability during long-time generation, yielding state-of-the-art performance compared to existing methods.
title EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.16726