Revisiting Common Assumptions about Arabic Dialects in NLP

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Keleg, Amr, Goldwater, Sharon, Magdy, Walid
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912399363145728
author Keleg, Amr
Goldwater, Sharon
Magdy, Walid
author_facet Keleg, Amr
Goldwater, Sharon
Magdy, Walid
contents Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g., ``Arabic dialects can be grouped into distinguishable regional dialects") and are manifested in different computational tasks such as Arabic Dialect Identification (ADI). However, these assumptions are not quantitatively verified. We identify four of these assumptions and examine them by extending and analyzing a multi-label dataset, where the validity of each sentence in 11 different country-level dialects is manually assessed by speakers of these dialects. Our analysis indicates that the four assumptions oversimplify reality, and some of them are not always accurate. This in turn might be hindering further progress in different Arabic NLP tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21816
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Common Assumptions about Arabic Dialects in NLP
Keleg, Amr
Goldwater, Sharon
Magdy, Walid
Computation and Language
Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g., ``Arabic dialects can be grouped into distinguishable regional dialects") and are manifested in different computational tasks such as Arabic Dialect Identification (ADI). However, these assumptions are not quantitatively verified. We identify four of these assumptions and examine them by extending and analyzing a multi-label dataset, where the validity of each sentence in 11 different country-level dialects is manually assessed by speakers of these dialects. Our analysis indicates that the four assumptions oversimplify reality, and some of them are not always accurate. This in turn might be hindering further progress in different Arabic NLP tasks.
title Revisiting Common Assumptions about Arabic Dialects in NLP
topic Computation and Language
url https://arxiv.org/abs/2505.21816