NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Talafha, Bashar, Toyin, Hawau Olamide, Sullivan, Peter, Elmadany, AbdelRahim, Juma, Abdurrahman, Djanibekov, Amirbek, Zhang, Chiyu, Alshehhi, Hamad, Aldarmaki, Hanan, Jarrar, Mustafa, Habash, Nizar, Abdul-Mageed, Muhammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909769181167616
author Talafha, Bashar
Toyin, Hawau Olamide
Sullivan, Peter
Elmadany, AbdelRahim
Juma, Abdurrahman
Djanibekov, Amirbek
Zhang, Chiyu
Alshehhi, Hamad
Aldarmaki, Hanan
Jarrar, Mustafa
Habash, Nizar
Abdul-Mageed, Muhammad
author_facet Talafha, Bashar
Toyin, Hawau Olamide
Sullivan, Peter
Elmadany, AbdelRahim
Juma, Abdurrahman
Djanibekov, Amirbek
Zhang, Chiyu
Alshehhi, Hamad
Aldarmaki, Hanan
Jarrar, Mustafa
Habash, Nizar
Abdul-Mageed, Muhammad
contents We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken dialect identification (Subtask 1), speech recognition (Subtask 2), and diacritic restoration for spoken dialects (Subtask 3). A total of 44 teams registered, and during the testing phase, 100 valid submissions were received from eight unique teams. The distribution was as follows: 34 submissions for Subtask 1 "five teamsæ, 47 submissions for Subtask 2 "six teams", and 19 submissions for Subtask 3 "two teams". The best-performing systems achieved 79.8% accuracy on Subtask 1, 35.68/12.20 WER/CER (overall average) on Subtask 2, and 55/13 WER/CER on Subtask 3. These results highlight the ongoing challenges of Arabic dialect speech processing, particularly in dialect identification, recognition, and diacritic restoration. We also summarize the methods adopted by participating teams and briefly outline directions for future editions of NADI.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
Talafha, Bashar
Toyin, Hawau Olamide
Sullivan, Peter
Elmadany, AbdelRahim
Juma, Abdurrahman
Djanibekov, Amirbek
Zhang, Chiyu
Alshehhi, Hamad
Aldarmaki, Hanan
Jarrar, Mustafa
Habash, Nizar
Abdul-Mageed, Muhammad
Computation and Language
Sound
We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken dialect identification (Subtask 1), speech recognition (Subtask 2), and diacritic restoration for spoken dialects (Subtask 3). A total of 44 teams registered, and during the testing phase, 100 valid submissions were received from eight unique teams. The distribution was as follows: 34 submissions for Subtask 1 "five teamsæ, 47 submissions for Subtask 2 "six teams", and 19 submissions for Subtask 3 "two teams". The best-performing systems achieved 79.8% accuracy on Subtask 1, 35.68/12.20 WER/CER (overall average) on Subtask 2, and 55/13 WER/CER on Subtask 3. These results highlight the ongoing challenges of Arabic dialect speech processing, particularly in dialect identification, recognition, and diacritic restoration. We also summarize the methods adopted by participating teams and briefly outline directions for future editions of NADI.
title NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
topic Computation and Language
Sound
url https://arxiv.org/abs/2509.02038