ADI-20: Arabic Dialect Identification dataset and models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Elleuch, Haroun, Mdhaffar, Salima, Estève, Yannick, Bougares, Fethi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917078306390016
author Elleuch, Haroun
Mdhaffar, Salima
Estève, Yannick
Bougares, Fethi
author_facet Elleuch, Haroun
Mdhaffar, Salima
Estève, Yannick
Bougares, Fethi
contents We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises 3,556 hours from 19 Arabic dialects in addition to Modern Standard Arabic (MSA). We used this dataset to train and evaluate various state-of-the-art ADI systems. We explored fine-tuning pre-trained ECAPA-TDNN-based models, as well as Whisper encoder blocks coupled with an attention pooling layer and a classification dense layer. We investigated the effect of (i) training data size and (ii) the model's number of parameters on identification performance. Our results show a small decrease in F1 score while using only 30% of the original training data. We open-source our collected data and trained models to enable the reproduction of our work, as well as support further research in ADI.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ADI-20: Arabic Dialect Identification dataset and models
Elleuch, Haroun
Mdhaffar, Salima
Estève, Yannick
Bougares, Fethi
Computation and Language
We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises 3,556 hours from 19 Arabic dialects in addition to Modern Standard Arabic (MSA). We used this dataset to train and evaluate various state-of-the-art ADI systems. We explored fine-tuning pre-trained ECAPA-TDNN-based models, as well as Whisper encoder blocks coupled with an attention pooling layer and a classification dense layer. We investigated the effect of (i) training data size and (ii) the model's number of parameters on identification performance. Our results show a small decrease in F1 score while using only 30% of the original training data. We open-source our collected data and trained models to enable the reproduction of our work, as well as support further research in ADI.
title ADI-20: Arabic Dialect Identification dataset and models
topic Computation and Language
url https://arxiv.org/abs/2511.10070