MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Al-Henaki, Lubna, Al-Khalifa, Hend, Al-Salman, Abdulmalik, Alqubayshi, Hajar, Al-Twailay, Hind, Alghamdi, Gheeda, Aljasim, Hawra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913700252745728
author Al-Henaki, Lubna
Al-Khalifa, Hend
Al-Salman, Abdulmalik
Alqubayshi, Hajar
Al-Twailay, Hind
Alghamdi, Gheeda
Aljasim, Hawra
author_facet Al-Henaki, Lubna
Al-Khalifa, Hend
Al-Salman, Abdulmalik
Alqubayshi, Hajar
Al-Twailay, Hind
Alghamdi, Gheeda
Aljasim, Hawra
contents Propaganda is a form of persuasion that has been used throughout history with the intention goal of influencing people's opinions through rhetorical and psychological persuasion techniques for determined ends. Although Arabic ranked as the fourth most-used language on the internet, resources for propaganda detection in languages other than English, especially Arabic, remain extremely limited. To address this gap, the first Arabic dataset for Multi-label Propaganda, Sentiment, and Emotion (MultiProSE) has been introduced. MultiProSE is an open-source extension of the existing Arabic propaganda dataset, ArPro, with the addition of sentiment and emotion annotations for each text. This dataset comprises 8,000 annotated news articles, which is the largest propaganda dataset to date. For each task, several baselines have been developed using large language models (LLMs), such as GPT-4o-mini, and pre-trained language models (PLMs), including three BERT-based models. The dataset, annotation guidelines, and source code are all publicly released to facilitate future research and development in Arabic language models and contribute to a deeper understanding of how various opinion dimensions interact in news media1.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08319
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection
Al-Henaki, Lubna
Al-Khalifa, Hend
Al-Salman, Abdulmalik
Alqubayshi, Hajar
Al-Twailay, Hind
Alghamdi, Gheeda
Aljasim, Hawra
Computation and Language
Propaganda is a form of persuasion that has been used throughout history with the intention goal of influencing people's opinions through rhetorical and psychological persuasion techniques for determined ends. Although Arabic ranked as the fourth most-used language on the internet, resources for propaganda detection in languages other than English, especially Arabic, remain extremely limited. To address this gap, the first Arabic dataset for Multi-label Propaganda, Sentiment, and Emotion (MultiProSE) has been introduced. MultiProSE is an open-source extension of the existing Arabic propaganda dataset, ArPro, with the addition of sentiment and emotion annotations for each text. This dataset comprises 8,000 annotated news articles, which is the largest propaganda dataset to date. For each task, several baselines have been developed using large language models (LLMs), such as GPT-4o-mini, and pre-trained language models (PLMs), including three BERT-based models. The dataset, annotation guidelines, and source code are all publicly released to facilitate future research and development in Arabic language models and contribute to a deeper understanding of how various opinion dimensions interact in news media1.
title MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection
topic Computation and Language
url https://arxiv.org/abs/2502.08319