ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xulong, Qu, Xiaoyang, Shi, Haoxiang, Xiao, Chunguang, Wang, Jianzong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910713964920832
author Zhang, Xulong
Qu, Xiaoyang
Shi, Haoxiang
Xiao, Chunguang
Wang, Jianzong
author_facet Zhang, Xulong
Qu, Xiaoyang
Shi, Haoxiang
Xiao, Chunguang
Wang, Jianzong
contents This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations that lack emotional depth and variety, failing to align with human expectations. To overcome these limitations, we introduce a novel STA model coupled with a reward model. This combination enables the decoupling of emotion and content under audio conditions through a cross-coupling training approach. Additionally, we develop a training methodology that leverages automatic quality evaluation of generated facial animations to guide the reinforcement learning process. This methodology encourages the STA model to explore a broader range of possibilities, resulting in the generation of diverse and emotionally expressive facial animations of superior quality. We conduct extensive empirical experiments on a benchmark dataset, and the results validate the effectiveness of our proposed framework in generating high-quality, emotionally rich 3D animations that are better aligned with human preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13089
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
Zhang, Xulong
Qu, Xiaoyang
Shi, Haoxiang
Xiao, Chunguang
Wang, Jianzong
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations that lack emotional depth and variety, failing to align with human expectations. To overcome these limitations, we introduce a novel STA model coupled with a reward model. This combination enables the decoupling of emotion and content under audio conditions through a cross-coupling training approach. Additionally, we develop a training methodology that leverages automatic quality evaluation of generated facial animations to guide the reinforcement learning process. This methodology encourages the STA model to explore a broader range of possibilities, resulting in the generation of diverse and emotionally expressive facial animations of superior quality. We conduct extensive empirical experiments on a benchmark dataset, and the results validate the effectiveness of our proposed framework in generating high-quality, emotionally rich 3D animations that are better aligned with human preferences.
title ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
topic Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.13089