Towards Zero-Shot Text-To-Speech for Arabic Dialects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Doan, Khai Duy, Waheed, Abdul, Abdul-Mageed, Muhammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916315204157440
author Doan, Khai Duy
Waheed, Abdul
Abdul-Mageed, Muhammad
author_facet Doan, Khai Duy
Waheed, Abdul
Abdul-Mageed, Muhammad
contents Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native speakers, by first adapting a sizeable existing dataset to suit the needs of speech synthesis. Additionally, we employ a set of Arabic dialect identification models to explore the impact of pre-defined dialect labels on improving the ZS-TTS model in a multi-dialect setting. Subsequently, we fine-tune the XTTS\footnote{https://docs.coqui.ai/en/latest/models/xtts.html}\footnote{https://medium.com/machine-learns/xtts-v2-new-version-of-the-open-source-text-to-speech-model-af73914db81f}\footnote{https://medium.com/@erogol/xtts-v1-techincal-notes-eb83ff05bdc} model, an open-source architecture. We then evaluate our models on a dataset comprising 31 unseen speakers and an in-house dialectal dataset. Our automated and human evaluation results show convincing performance while capable of generating dialectal speech. Our study highlights significant potential for improvements in this emerging area of research in Arabic.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Zero-Shot Text-To-Speech for Arabic Dialects
Doan, Khai Duy
Waheed, Abdul
Abdul-Mageed, Muhammad
Computation and Language
Sound
Audio and Speech Processing
Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native speakers, by first adapting a sizeable existing dataset to suit the needs of speech synthesis. Additionally, we employ a set of Arabic dialect identification models to explore the impact of pre-defined dialect labels on improving the ZS-TTS model in a multi-dialect setting. Subsequently, we fine-tune the XTTS\footnote{https://docs.coqui.ai/en/latest/models/xtts.html}\footnote{https://medium.com/machine-learns/xtts-v2-new-version-of-the-open-source-text-to-speech-model-af73914db81f}\footnote{https://medium.com/@erogol/xtts-v1-techincal-notes-eb83ff05bdc} model, an open-source architecture. We then evaluate our models on a dataset comprising 31 unseen speakers and an in-house dialectal dataset. Our automated and human evaluation results show convincing performance while capable of generating dialectal speech. Our study highlights significant potential for improvements in this emerging area of research in Arabic.
title Towards Zero-Shot Text-To-Speech for Arabic Dialects
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.16751