KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913308304474112 |
|---|---|
| author | Abilbekov, Adal Mussakhojayeva, Saida Yeshpanov, Rustem Varol, Huseyin Atakan |
| author_facet | Abilbekov, Adal Mussakhojayeva, Saida Yeshpanov, Rustem Varol, Huseyin Atakan |
| contents | This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male narrators. The list of the emotions considered include "neutral", "angry", "happy", "sad", "scared", and "surprised". We also developed a TTS model trained on the KazEmoTTS dataset. Objective and subjective evaluations were employed to assess the quality of synthesized speech, yielding an MCD score within the range of 6.02 to 7.67, alongside a MOS that spanned from 3.51 to 3.57. To facilitate reproducibility and inspire further research, we have made our code, pre-trained model, and dataset accessible in our GitHub repository. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_01033 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis Abilbekov, Adal Mussakhojayeva, Saida Yeshpanov, Rustem Varol, Huseyin Atakan Audio and Speech Processing This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male narrators. The list of the emotions considered include "neutral", "angry", "happy", "sad", "scared", and "surprised". We also developed a TTS model trained on the KazEmoTTS dataset. Objective and subjective evaluations were employed to assess the quality of synthesized speech, yielding an MCD score within the range of 6.02 to 7.67, alongside a MOS that spanned from 3.51 to 3.57. To facilitate reproducibility and inspire further research, we have made our code, pre-trained model, and dataset accessible in our GitHub repository. |
| title | KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2404.01033 |