Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jingyi, Le, Hieu, Shu, Zhixin, Wang, Yang, Tsai, Yi-Hsuan, Samaras, Dimitris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909328824336384
author Xu, Jingyi
Le, Hieu
Shu, Zhixin
Wang, Yang
Tsai, Yi-Hsuan
Samaras, Dimitris
author_facet Xu, Jingyi
Le, Hieu
Shu, Zhixin
Wang, Yang
Tsai, Yi-Hsuan
Samaras, Dimitris
contents Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by previous audio-driven talking-head generation methods, which often results in static emotional outputs. In this paper, we explore how emotion intensity fluctuates during speech, proposing a method for capturing and generating these subtle shifts for talking-head generation. Specifically, we develop a talking-head framework that is capable of generating a variety of emotions with precise control over intensity levels. This is achieved by learning a continuous emotion latent space, where emotion types are encoded within latent orientations and emotion intensity is reflected in latent norms. In addition, to capture the dynamic intensity fluctuations, we adopt an audio-to-intensity predictor by considering the speaking tone that reflects the intensity. The training signals for this predictor are obtained through our emotion-agnostic intensity pseudo-labeling method without the need of frame-wise intensity labeling. Extensive experiments and analyses validate the effectiveness of our proposed method in accurately capturing and reproducing emotion intensity fluctuations in talking-head generation, thereby significantly enhancing the expressiveness and realism of the generated outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19501
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
Xu, Jingyi
Le, Hieu
Shu, Zhixin
Wang, Yang
Tsai, Yi-Hsuan
Samaras, Dimitris
Sound
Artificial Intelligence
Audio and Speech Processing
Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by previous audio-driven talking-head generation methods, which often results in static emotional outputs. In this paper, we explore how emotion intensity fluctuates during speech, proposing a method for capturing and generating these subtle shifts for talking-head generation. Specifically, we develop a talking-head framework that is capable of generating a variety of emotions with precise control over intensity levels. This is achieved by learning a continuous emotion latent space, where emotion types are encoded within latent orientations and emotion intensity is reflected in latent norms. In addition, to capture the dynamic intensity fluctuations, we adopt an audio-to-intensity predictor by considering the speaking tone that reflects the intensity. The training signals for this predictor are obtained through our emotion-agnostic intensity pseudo-labeling method without the need of frame-wise intensity labeling. Extensive experiments and analyses validate the effectiveness of our proposed method in accurately capturing and reproducing emotion intensity fluctuations in talking-head generation, thereby significantly enhancing the expressiveness and realism of the generated outputs.
title Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2409.19501