More Data, Fewer Diacritics: Scaling Arabic TTS

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Musleh, Ahmed, Zhang, Yifan, Darwish, Kareem
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911478563471360
author Musleh, Ahmed
Zhang, Yifan
Darwish, Kareem
author_facet Musleh, Ahmed
Zhang, Yifan
Darwish, Kareem
contents Arabic Text-to-Speech (TTS) research has been hindered by the availability of both publicly available training data and accurate Arabic diacritization models. In this paper, we address the limitation by exploring Arabic TTS training on large automatically annotated data. Namely, we built a robust pipeline for collecting Arabic recordings and processing them automatically using voice activity detection, speech recognition, automatic diacritization, and noise filtering, resulting in around 4,000 hours of Arabic TTS training data. We then trained several robust TTS models with voice cloning using varying amounts of data, namely 100, 1,000, and 4,000 hours with and without diacritization. We show that though models trained on diacritized data are generally better, larger amounts of training data compensate for the lack of diacritics to a significant degree. We plan to release a public Arabic TTS model that works without the need for diacritization.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01622
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle More Data, Fewer Diacritics: Scaling Arabic TTS
Musleh, Ahmed
Zhang, Yifan
Darwish, Kareem
Computation and Language
Arabic Text-to-Speech (TTS) research has been hindered by the availability of both publicly available training data and accurate Arabic diacritization models. In this paper, we address the limitation by exploring Arabic TTS training on large automatically annotated data. Namely, we built a robust pipeline for collecting Arabic recordings and processing them automatically using voice activity detection, speech recognition, automatic diacritization, and noise filtering, resulting in around 4,000 hours of Arabic TTS training data. We then trained several robust TTS models with voice cloning using varying amounts of data, namely 100, 1,000, and 4,000 hours with and without diacritization. We show that though models trained on diacritized data are generally better, larger amounts of training data compensate for the lack of diacritics to a significant degree. We plan to release a public Arabic TTS model that works without the need for diacritization.
title More Data, Fewer Diacritics: Scaling Arabic TTS
topic Computation and Language
url https://arxiv.org/abs/2603.01622