MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Elozeiri, Kareem, Abassy, Mervat, Nakov, Preslav, Wang, Yuxia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912813330464768
author Elozeiri, Kareem
Abassy, Mervat
Nakov, Preslav
Wang, Yuxia
author_facet Elozeiri, Kareem
Abassy, Mervat
Nakov, Preslav
Wang, Yuxia
contents Commonsense validation evaluates whether a sentence aligns with everyday human understanding, a critical capability for developing robust natural language understanding systems. While substantial progress has been made in English, the task remains underexplored in Arabic, particularly given its rich linguistic diversity. Existing Arabic resources have primarily focused on Modern Standard Arabic (MSA), leaving regional dialects underrepresented despite their prevalence in spoken contexts. To bridge this gap, we present two key contributions. We introduce MuDRiC, an extended Arabic commonsense dataset incorporating multiple dialects. To the best of our knowledge, this is the first Arabic multi-dialect commonsense reasoning dataset. We further propose a novel method adapting Graph Convolutional Networks (GCNs) to Arabic commonsense reasoning, which enhances semantic relationship modeling for improved commonsense validation. Our experimental results demonstrate that this approach consistently outperforms the baseline of direct language model fine-tuning. Overall, our work enhances Arabic natural language understanding by providing a foundational dataset and a new method for handling its complex variations. Data and code are available at https://github.com/KareemElozeiri/MuDRiC.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13130
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
Elozeiri, Kareem
Abassy, Mervat
Nakov, Preslav
Wang, Yuxia
Computation and Language
Commonsense validation evaluates whether a sentence aligns with everyday human understanding, a critical capability for developing robust natural language understanding systems. While substantial progress has been made in English, the task remains underexplored in Arabic, particularly given its rich linguistic diversity. Existing Arabic resources have primarily focused on Modern Standard Arabic (MSA), leaving regional dialects underrepresented despite their prevalence in spoken contexts. To bridge this gap, we present two key contributions. We introduce MuDRiC, an extended Arabic commonsense dataset incorporating multiple dialects. To the best of our knowledge, this is the first Arabic multi-dialect commonsense reasoning dataset. We further propose a novel method adapting Graph Convolutional Networks (GCNs) to Arabic commonsense reasoning, which enhances semantic relationship modeling for improved commonsense validation. Our experimental results demonstrate that this approach consistently outperforms the baseline of direct language model fine-tuning. Overall, our work enhances Arabic natural language understanding by providing a foundational dataset and a new method for handling its complex variations. Data and code are available at https://github.com/KareemElozeiri/MuDRiC.
title MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
topic Computation and Language
url https://arxiv.org/abs/2508.13130