Fine-Grained Quantitative Emotion Editing for Speech Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Inoue, Sho, Zhou, Kun, Wang, Shuai, Li, Haizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913521536598016
author Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
author_facet Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
contents It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulating the rendering of emotions for speech generation. We propose a hierarchical emotion distribution extractor, i.e. Hierarchical ED, that quantifies the intensity of emotions at different levels of granularity. Support vector machines (SVMs) are employed to rank emotion intensity, resulting in a hierarchical emotional embedding. Hierarchical ED is subsequently integrated into the FastSpeech2 framework, guiding the model to learn emotion intensity at phoneme, word, and utterance levels. During synthesis, users can manually edit the emotional intensity of the generated voices. Both objective and subjective evaluations demonstrate the effectiveness of the proposed network in terms of fine-grained quantitative emotion editing.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02002
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-Grained Quantitative Emotion Editing for Speech Generation
Inoue, Sho
Zhou, Kun
Wang, Shuai
Li, Haizhou
Sound
Audio and Speech Processing
It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulating the rendering of emotions for speech generation. We propose a hierarchical emotion distribution extractor, i.e. Hierarchical ED, that quantifies the intensity of emotions at different levels of granularity. Support vector machines (SVMs) are employed to rank emotion intensity, resulting in a hierarchical emotional embedding. Hierarchical ED is subsequently integrated into the FastSpeech2 framework, guiding the model to learn emotion intensity at phoneme, word, and utterance levels. During synthesis, users can manually edit the emotional intensity of the generated voices. Both objective and subjective evaluations demonstrate the effectiveness of the proposed network in terms of fine-grained quantitative emotion editing.
title Fine-Grained Quantitative Emotion Editing for Speech Generation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2403.02002