Hierarchical Control of Emotion Rendering in Speech Synthesis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Inoue, Sho, Zhou, Kun, Wang, Shuai, Li, Haizhou
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912443529166848
author Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
author_facet Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
contents Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based emotional TTS framework with a novel approach for emotion intensity modeling to facilitate fine-grained control over emotion rendering at the phoneme, word, and utterance levels. We introduce a hierarchical emotion distribution (ED) extractor that captures a quantifiable ED embedding across different speech segment levels. Additionally, we explore various acoustic features and assess their impact on emotion intensity modeling. During TTS training, the hierarchical ED embedding effectively captures the variance in emotion intensity from the reference audio and correlates it with linguistic and speaker information. The TTS model not only generates emotional speech during inference, but also quantitatively controls the emotion rendering over the speech constituents. Both objective and subjective evaluations demonstrate the effectiveness of our framework in terms of speech quality, emotional expressiveness, and hierarchical emotion control.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12498
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hierarchical Control of Emotion Rendering in Speech Synthesis
Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
Sound
Audio and Speech Processing
Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based emotional TTS framework with a novel approach for emotion intensity modeling to facilitate fine-grained control over emotion rendering at the phoneme, word, and utterance levels. We introduce a hierarchical emotion distribution (ED) extractor that captures a quantifiable ED embedding across different speech segment levels. Additionally, we explore various acoustic features and assess their impact on emotion intensity modeling. During TTS training, the hierarchical ED embedding effectively captures the variance in emotion intensity from the reference audio and correlates it with linguistic and speaker information. The TTS model not only generates emotional speech during inference, but also quantitatively controls the emotion rendering over the speech constituents. Both objective and subjective evaluations demonstrate the effectiveness of our framework in terms of speech quality, emotional expressiveness, and hierarchical emotion control.
title Hierarchical Control of Emotion Rendering in Speech Synthesis
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.12498