EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zaghouani, Wajdi, Biswas, Md. Rafiul
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916747488002048
author Zaghouani, Wajdi
Biswas, Md. Rafiul
author_facet Zaghouani, Wajdi
Biswas, Md. Rafiul
contents This research introduces a bilingual dataset comprising 23,456 entries for Arabic and 10,036 entries for English, annotated for emotions and hope speech, addressing the scarcity of multi-emotion (Emotion and hope) datasets. The dataset provides comprehensive annotations capturing emotion intensity, complexity, and causes, alongside detailed classifications and subcategories for hope speech. To ensure annotation reliability, Fleiss' Kappa was employed, revealing 0.75-0.85 agreement among annotators both for Arabic and English language. The evaluation metrics (micro-F1-Score=0.67) obtained from the baseline model (i.e., using a machine learning model) validate that the data annotations are worthy. This dataset offers a valuable resource for advancing natural language processing in underrepresented languages, fostering better cross-linguistic analysis of emotions and hope speech.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11959
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic
Zaghouani, Wajdi
Biswas, Md. Rafiul
Computation and Language
This research introduces a bilingual dataset comprising 23,456 entries for Arabic and 10,036 entries for English, annotated for emotions and hope speech, addressing the scarcity of multi-emotion (Emotion and hope) datasets. The dataset provides comprehensive annotations capturing emotion intensity, complexity, and causes, alongside detailed classifications and subcategories for hope speech. To ensure annotation reliability, Fleiss' Kappa was employed, revealing 0.75-0.85 agreement among annotators both for Arabic and English language. The evaluation metrics (micro-F1-Score=0.67) obtained from the baseline model (i.e., using a machine learning model) validate that the data annotations are worthy. This dataset offers a valuable resource for advancing natural language processing in underrepresented languages, fostering better cross-linguistic analysis of emotions and hope speech.
title EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic
topic Computation and Language
url https://arxiv.org/abs/2505.11959