Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Inoue, Sho, Zhou, Kun, Wang, Shuai, Li, Haizhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910448218013696
author Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
author_facet Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
contents It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with linguistic prosody. Our goal is to construct a hierarchical emotion distribution (ED) that effectively encapsulates intensity variations of emotions at various levels of granularity, encompassing phonemes, words, and utterances. During TTS training, the hierarchical ED is extracted from the ground-truth audio and guides the predictor to establish a connection between emotional and linguistic prosody. At run-time inference, the TTS model generates emotional speech and, at the same time, provides quantitative control of emotion over the speech constituents. Both objective and subjective evaluations validate the effectiveness of the proposed framework in terms of emotion prediction and control.
format Preprint
id arxiv_https___arxiv_org_abs_2405_09171
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
Sound
Audio and Speech Processing
It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with linguistic prosody. Our goal is to construct a hierarchical emotion distribution (ED) that effectively encapsulates intensity variations of emotions at various levels of granularity, encompassing phonemes, words, and utterances. During TTS training, the hierarchical ED is extracted from the ground-truth audio and guides the predictor to establish a connection between emotional and linguistic prosody. At run-time inference, the TTS model generates emotional speech and, at the same time, provides quantitative control of emotion over the speech constituents. Both objective and subjective evaluations validate the effectiveness of the proposed framework in terms of emotion prediction and control.
title Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.09171