Exploiting Dialect Identification in Automatic Dialectal Text Normalization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Alhafni, Bashar, Al-Towaity, Sarah, Fawzy, Ziyad, Nassar, Fatema, Eryani, Fadhl, Bouamor, Houda, Habash, Nizar
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916766680088576
author Alhafni, Bashar
Al-Towaity, Sarah
Fawzy, Ziyad
Nassar, Fatema
Eryani, Fadhl
Bouamor, Houda
Habash, Nizar
author_facet Alhafni, Bashar
Al-Towaity, Sarah
Fawzy, Ziyad
Nassar, Fatema
Eryani, Fadhl
Bouamor, Houda
Habash, Nizar
contents Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do not have standard orthographies. This, combined with the inherent noise in user-generated content on social media, presents a major challenge to NLP applications dealing with Dialectal Arabic. In this paper, we explore and report on the task of CODAfication, which aims to normalize Dialectal Arabic into the Conventional Orthography for Dialectal Arabic (CODA). We work with a unique parallel corpus of multiple Arabic dialects focusing on five major city dialects. We benchmark newly developed pretrained sequence-to-sequence models on the task of CODAfication. We further show that using dialect identification information improves the performance across all dialects. We make our code, data, and pretrained models publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03020
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting Dialect Identification in Automatic Dialectal Text Normalization
Alhafni, Bashar
Al-Towaity, Sarah
Fawzy, Ziyad
Nassar, Fatema
Eryani, Fadhl
Bouamor, Houda
Habash, Nizar
Computation and Language
Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do not have standard orthographies. This, combined with the inherent noise in user-generated content on social media, presents a major challenge to NLP applications dealing with Dialectal Arabic. In this paper, we explore and report on the task of CODAfication, which aims to normalize Dialectal Arabic into the Conventional Orthography for Dialectal Arabic (CODA). We work with a unique parallel corpus of multiple Arabic dialects focusing on five major city dialects. We benchmark newly developed pretrained sequence-to-sequence models on the task of CODAfication. We further show that using dialect identification information improves the performance across all dialects. We make our code, data, and pretrained models publicly available.
title Exploiting Dialect Identification in Automatic Dialectal Text Normalization
topic Computation and Language
url https://arxiv.org/abs/2407.03020