Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Weipeng, Lin, Chuming, Xu, Chengming, Xu, FeiFan, Hu, Xiaobin, Ji, Xiaozhong, Zhu, Junwei, Wang, Chengjie, Fu, Yanwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908480005210112
author Tan, Weipeng
Lin, Chuming
Xu, Chengming
Xu, FeiFan
Hu, Xiaobin
Ji, Xiaozhong
Zhu, Junwei
Wang, Chengjie
Fu, Yanwei
author_facet Tan, Weipeng
Lin, Chuming
Xu, Chengming
Xu, FeiFan
Hu, Xiaobin
Ji, Xiaozhong
Zhu, Junwei
Wang, Chengjie
Fu, Yanwei
contents Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: insufficient utilization of audio's inherent emotional cues, identity leakage in emotion representations, and isolated learning of emotion correlations. To address these challenges, we propose a novel framework dubbed as DICE-Talk, following the idea of disentangling identity with emotion, and then cooperating emotions with similar characteristics. First, we develop a disentangled emotion embedder that jointly models audio-visual emotional cues through cross-modal attention, representing emotions as identity-agnostic Gaussian distributions. Second, we introduce a correlation-enhanced emotion conditioning module with learnable Emotion Banks that explicitly capture inter-emotion relationships through vector quantization and attention-based feature aggregation. Third, we design an emotion discrimination objective that enforces affective consistency during the diffusion process through latent-space classification. Extensive experiments on MEAD and HDTF datasets demonstrate our method's superiority, outperforming state-of-the-art approaches in emotion accuracy while maintaining competitive lip-sync performance. Qualitative results and user studies further confirm our method's ability to generate identity-preserving portraits with rich, correlated emotional expressions that naturally adapt to unseen identities.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
Tan, Weipeng
Lin, Chuming
Xu, Chengming
Xu, FeiFan
Hu, Xiaobin
Ji, Xiaozhong
Zhu, Junwei
Wang, Chengjie
Fu, Yanwei
Computer Vision and Pattern Recognition
Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: insufficient utilization of audio's inherent emotional cues, identity leakage in emotion representations, and isolated learning of emotion correlations. To address these challenges, we propose a novel framework dubbed as DICE-Talk, following the idea of disentangling identity with emotion, and then cooperating emotions with similar characteristics. First, we develop a disentangled emotion embedder that jointly models audio-visual emotional cues through cross-modal attention, representing emotions as identity-agnostic Gaussian distributions. Second, we introduce a correlation-enhanced emotion conditioning module with learnable Emotion Banks that explicitly capture inter-emotion relationships through vector quantization and attention-based feature aggregation. Third, we design an emotion discrimination objective that enforces affective consistency during the diffusion process through latent-space classification. Extensive experiments on MEAD and HDTF datasets demonstrate our method's superiority, outperforming state-of-the-art approaches in emotion accuracy while maintaining competitive lip-sync performance. Qualitative results and user studies further confirm our method's ability to generate identity-preserving portraits with rich, correlated emotional expressions that naturally adapt to unseen identities.
title Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.18087