ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910970322878464 |
|---|---|
| author | Toyin, Hawau Olamide Marew, Rufael Alblooshi, Humaid Magdy, Samar M. Aldarmaki, Hanan |
| author_facet | Toyin, Hawau Olamide Marew, Rufael Alblooshi, Humaid Magdy, Samar M. Aldarmaki, Hanan |
| contents | We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_20506 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Toyin, Hawau Olamide Marew, Rufael Alblooshi, Humaid Magdy, Samar M. Aldarmaki, Hanan Computation and Language Artificial Intelligence Sound Audio and Speech Processing We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use. |
| title | ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis |
| topic | Computation and Language Artificial Intelligence Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.20506 |