ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910385053892608 |
|---|---|
| author | Ogunremi, Tolulope Tubosun, Kola Aremu, Anuoluwapo Orife, Iroro Adelani, David Ifeoluwa |
| author_facet | Ogunremi, Tolulope Tubosun, Kola Aremu, Anuoluwapo Orife, Iroro Adelani, David Ifeoluwa |
| contents | We introduce ÌròyìnSpeech, a new corpus influenced by the desire to increase the amount of high quality, contemporary Yorùbá speech data, which can be used for both Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) tasks. We curated about 23000 text sentences from news and creative writing domains with the open license CC-BY-4.0. To encourage a participatory approach to data creation, we provide 5000 curated sentences to the Mozilla Common Voice platform to crowd-source the recording and validation of Yorùbá speech data. In total, we created about 42 hours of speech data recorded by 80 volunteers in-house, and 6 hours of validated recordings on Mozilla Common Voice platform. Our TTS evaluation suggests that a high-fidelity, general domain, single-speaker Yorùbá voice is possible with as little as 5 hours of speech. Similarly, for ASR we obtained a baseline word error rate (WER) of 23.8. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2307_16071 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus Ogunremi, Tolulope Tubosun, Kola Aremu, Anuoluwapo Orife, Iroro Adelani, David Ifeoluwa Computation and Language Sound Audio and Speech Processing We introduce ÌròyìnSpeech, a new corpus influenced by the desire to increase the amount of high quality, contemporary Yorùbá speech data, which can be used for both Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) tasks. We curated about 23000 text sentences from news and creative writing domains with the open license CC-BY-4.0. To encourage a participatory approach to data creation, we provide 5000 curated sentences to the Mozilla Common Voice platform to crowd-source the recording and validation of Yorùbá speech data. In total, we created about 42 hours of speech data recorded by 80 volunteers in-house, and 6 hours of validated recordings on Mozilla Common Voice platform. Our TTS evaluation suggests that a high-fidelity, general domain, single-speaker Yorùbá voice is possible with as little as 5 hours of speech. Similarly, for ASR we obtained a baseline word error rate (WER) of 23.8. |
| title | ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2307.16071 |