EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cong, Gaoxiang, Pan, Jiadong, Li, Liang, Qi, Yuankai, Peng, Yuxin, Hengel, Anton van den, Yang, Jian, Huang, Qingming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912345926664192
author Cong, Gaoxiang
Pan, Jiadong
Li, Liang
Qi, Yuankai
Peng, Yuxin
Hengel, Anton van den
Yang, Jian
Huang, Qingming
author_facet Cong, Gaoxiang
Pan, Jiadong
Li, Liang
Qi, Yuankai
Peng, Yuxin
Hengel, Anton van den
Yang, Jian
Huang, Qingming
contents Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to simultaneously hold audio-visual sync and achieve clear pronunciation; (2) They lack the capacity to express user-defined emotions. To address these problems, we propose EmoDubber, an emotion-controllable dubbing architecture that allows users to specify emotion type and emotional intensity while satisfying high-quality lip sync and pronunciation. Specifically, we first design Lip-related Prosody Aligning (LPA), which focuses on learning the inherent consistency between lip motion and prosody variation by duration level contrastive learning to incorporate reasonable alignment. Then, we design Pronunciation Enhancing (PE) strategy to fuse the video-level phoneme sequences by efficient conformer to improve speech intelligibility. Next, the speaker identity adapting module aims to decode acoustics prior and inject the speaker style embedding. After that, the proposed Flow-based User Emotion Controlling (FUEC) is used to synthesize waveform by flow matching prediction network conditioned on acoustics prior. In this process, the FUEC determines the gradient direction and guidance scale based on the user's emotion instructions by the positive and negative guidance mechanism, which focuses on amplifying the desired emotion while suppressing others. Extensive experimental results on three benchmark datasets demonstrate favorable performance compared to several state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08988
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
Cong, Gaoxiang
Pan, Jiadong
Li, Liang
Qi, Yuankai
Peng, Yuxin
Hengel, Anton van den
Yang, Jian
Huang, Qingming
Sound
Multimedia
Audio and Speech Processing
Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to simultaneously hold audio-visual sync and achieve clear pronunciation; (2) They lack the capacity to express user-defined emotions. To address these problems, we propose EmoDubber, an emotion-controllable dubbing architecture that allows users to specify emotion type and emotional intensity while satisfying high-quality lip sync and pronunciation. Specifically, we first design Lip-related Prosody Aligning (LPA), which focuses on learning the inherent consistency between lip motion and prosody variation by duration level contrastive learning to incorporate reasonable alignment. Then, we design Pronunciation Enhancing (PE) strategy to fuse the video-level phoneme sequences by efficient conformer to improve speech intelligibility. Next, the speaker identity adapting module aims to decode acoustics prior and inject the speaker style embedding. After that, the proposed Flow-based User Emotion Controlling (FUEC) is used to synthesize waveform by flow matching prediction network conditioned on acoustics prior. In this process, the FUEC determines the gradient direction and guidance scale based on the user's emotion instructions by the positive and negative guidance mechanism, which focuses on amplifying the desired emotion while suppressing others. Extensive experimental results on three benchmark datasets demonstrate favorable performance compared to several state-of-the-art methods.
title EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2412.08988