Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdullah, Badr M., Baas, Matthew, Möbius, Bernd, Klakow, Dietrich
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913867816239104
author Abdullah, Badr M.
Baas, Matthew
Möbius, Bernd
Klakow, Dietrich
author_facet Abdullah, Badr M.
Baas, Matthew
Möbius, Bernd
Klakow, Dietrich
contents Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of current ADI systems is limited by poor generalization to out-of-domain speech. In this paper, we present an effective approach based on voice conversion for training ADI models that achieves state-of-the-art performance and significantly improves robustness in cross-domain scenarios. Evaluated on a newly collected real-world test set spanning four different domains, our approach yields consistent improvements of up to +34.1% in accuracy across domains. Furthermore, we present an analysis of our approach and demonstrate that voice conversion helps mitigate the speaker bias in the ADI dataset. We release our robust ADI model and cross-domain evaluation dataset to support the development of inclusive speech technologies for Arabic.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
Abdullah, Badr M.
Baas, Matthew
Möbius, Bernd
Klakow, Dietrich
Computation and Language
Sound
Audio and Speech Processing
Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of current ADI systems is limited by poor generalization to out-of-domain speech. In this paper, we present an effective approach based on voice conversion for training ADI models that achieves state-of-the-art performance and significantly improves robustness in cross-domain scenarios. Evaluated on a newly collected real-world test set spanning four different domains, our approach yields consistent improvements of up to +34.1% in accuracy across domains. Furthermore, we present an analysis of our approach and demonstrate that voice conversion helps mitigate the speaker bias in the ADI dataset. We release our robust ADI model and cross-domain evaluation dataset to support the development of inclusive speech technologies for Arabic.
title Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.24713