ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Toyin, Hawau Olamide, Marew, Rufael, Alblooshi, Humaid, Magdy, Samar M., Aldarmaki, Hanan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910970322878464
author Toyin, Hawau Olamide
Marew, Rufael
Alblooshi, Humaid
Magdy, Samar M.
Aldarmaki, Hanan
author_facet Toyin, Hawau Olamide
Marew, Rufael
Alblooshi, Humaid
Magdy, Samar M.
Aldarmaki, Hanan
contents We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
Toyin, Hawau Olamide
Marew, Rufael
Alblooshi, Humaid
Magdy, Samar M.
Aldarmaki, Hanan
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use.
title ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.20506