KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abilbekov, Adal, Mussakhojayeva, Saida, Yeshpanov, Rustem, Varol, Huseyin Atakan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913308304474112
author Abilbekov, Adal
Mussakhojayeva, Saida
Yeshpanov, Rustem
Varol, Huseyin Atakan
author_facet Abilbekov, Adal
Mussakhojayeva, Saida
Yeshpanov, Rustem
Varol, Huseyin Atakan
contents This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male narrators. The list of the emotions considered include "neutral", "angry", "happy", "sad", "scared", and "surprised". We also developed a TTS model trained on the KazEmoTTS dataset. Objective and subjective evaluations were employed to assess the quality of synthesized speech, yielding an MCD score within the range of 6.02 to 7.67, alongside a MOS that spanned from 3.51 to 3.57. To facilitate reproducibility and inspire further research, we have made our code, pre-trained model, and dataset accessible in our GitHub repository.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01033
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
Abilbekov, Adal
Mussakhojayeva, Saida
Yeshpanov, Rustem
Varol, Huseyin Atakan
Audio and Speech Processing
This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male narrators. The list of the emotions considered include "neutral", "angry", "happy", "sad", "scared", and "surprised". We also developed a TTS model trained on the KazEmoTTS dataset. Objective and subjective evaluations were employed to assess the quality of synthesized speech, yielding an MCD score within the range of 6.02 to 7.67, alongside a MOS that spanned from 3.51 to 3.57. To facilitate reproducibility and inspire further research, we have made our code, pre-trained model, and dataset accessible in our GitHub repository.
title KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
topic Audio and Speech Processing
url https://arxiv.org/abs/2404.01033