Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bellaouar, Slimane, Nehar, Attia, Souffi, Soumia, Bouameur, Mounia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915820758630400
author Bellaouar, Slimane
Nehar, Attia
Souffi, Soumia
Bouameur, Mounia
author_facet Bellaouar, Slimane
Nehar, Attia
Souffi, Soumia
Bouameur, Mounia
contents Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity analysis in Arabic. Recent advances in deep learning and Transformers have proven highly effective for text classification in English and French. This paper proposes a new approach for subjectivity assessment in Arabic textual data. To address the dearth of specialized annotated datasets, we developed a comprehensive dataset, AraDhati+, by leveraging existing Arabic datasets and collections (ASTD, LABR, HARD, and SANAD). Subsequently, we fine-tuned state-of-the-art Arabic language models (XLM-RoBERTa, AraBERT, and ArabianGPT) on AraDhati+ for effective subjectivity classification. Furthermore, we experimented with an ensemble decision approach to harness the strengths of individual models. Our approach achieves a remarkable accuracy of 97.79\,\% for Arabic subjectivity classification. Results demonstrate the effectiveness of the proposed approach in addressing the challenges posed by limited resources in Arabic language processing.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation
Bellaouar, Slimane
Nehar, Attia
Souffi, Soumia
Bouameur, Mounia
Computation and Language
Artificial Intelligence
Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity analysis in Arabic. Recent advances in deep learning and Transformers have proven highly effective for text classification in English and French. This paper proposes a new approach for subjectivity assessment in Arabic textual data. To address the dearth of specialized annotated datasets, we developed a comprehensive dataset, AraDhati+, by leveraging existing Arabic datasets and collections (ASTD, LABR, HARD, and SANAD). Subsequently, we fine-tuned state-of-the-art Arabic language models (XLM-RoBERTa, AraBERT, and ArabianGPT) on AraDhati+ for effective subjectivity classification. Furthermore, we experimented with an ensemble decision approach to harness the strengths of individual models. Our approach achieves a remarkable accuracy of 97.79\,\% for Arabic subjectivity classification. Results demonstrate the effectiveness of the proposed approach in addressing the challenges posed by limited resources in Arabic language processing.
title Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.19966