EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bian, Weizhen, Zhou, Yubo, Zhang, Kaitai, Gu, Xiaohan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909425180082176
author Bian, Weizhen
Zhou, Yubo
Zhang, Kaitai
Gu, Xiaohan
author_facet Bian, Weizhen
Zhou, Yubo
Zhang, Kaitai
Gu, Xiaohan
contents Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional expression, the development of TTS systems capable of controlling subtle emotional differences remains a formidable challenge. Existing emotional speech databases often suffer from overly simplistic labelling schemes that fail to capture a wide range of emotional states, thus limiting the effectiveness of emotion synthesis in TTS applications. To this end, recent efforts have focussed on building databases that use natural language annotations to describe speech emotions. However, these approaches are costly and require more emotional depth to train robust systems. In this paper, we propose a novel process aimed at building databases by systematically extracting emotion-rich speech segments and annotating them with detailed natural language descriptions through a generative model. This approach enhances the emotional granularity of the database and significantly reduces the reliance on costly manual annotations by automatically augmenting the data with high-level language models. The resulting rich database provides a scalable and economically viable solution for developing a more nuanced and dynamic basis for developing emotionally controlled TTS systems.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06581
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
Bian, Weizhen
Zhou, Yubo
Zhang, Kaitai
Gu, Xiaohan
Sound
Artificial Intelligence
Audio and Speech Processing
Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional expression, the development of TTS systems capable of controlling subtle emotional differences remains a formidable challenge. Existing emotional speech databases often suffer from overly simplistic labelling schemes that fail to capture a wide range of emotional states, thus limiting the effectiveness of emotion synthesis in TTS applications. To this end, recent efforts have focussed on building databases that use natural language annotations to describe speech emotions. However, these approaches are costly and require more emotional depth to train robust systems. In this paper, we propose a novel process aimed at building databases by systematically extracting emotion-rich speech segments and annotating them with detailed natural language descriptions through a generative model. This approach enhances the emotional granularity of the database and significantly reduces the reliance on costly manual annotations by automatically augmenting the data with high-level language models. The resulting rich database provides a scalable and economically viable solution for developing a more nuanced and dynamic basis for developing emotionally controlled TTS systems.
title EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.06581