EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dementieva, Daryna, Babakov, Nikolay, Fraser, Alexander
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911177453338624
author Dementieva, Daryna
Babakov, Nikolay
Fraser, Alexander
author_facet Dementieva, Daryna
Babakov, Nikolay
Fraser, Alexander
contents While Ukrainian NLP has seen progress in many texts processing tasks, emotion classification remains an underexplored area with no publicly available benchmark to date. In this work, we introduce EmoBench-UA, the first annotated dataset for emotion detection in Ukrainian texts. Our annotation schema is adapted from the previous English-centric works on emotion detection (Mohammad et al., 2018; Mohammad, 2022) guidelines. The dataset was created through crowdsourcing using the Toloka.ai platform ensuring high-quality of the annotation process. Then, we evaluate a range of approaches on the collected dataset, starting from linguistic-based baselines, synthetic data translated from English, to large language models (LLMs). Our findings highlight the challenges of emotion classification in non-mainstream languages like Ukrainian and emphasize the need for further development of Ukrainian-specific models and training resources.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian
Dementieva, Daryna
Babakov, Nikolay
Fraser, Alexander
Computation and Language
While Ukrainian NLP has seen progress in many texts processing tasks, emotion classification remains an underexplored area with no publicly available benchmark to date. In this work, we introduce EmoBench-UA, the first annotated dataset for emotion detection in Ukrainian texts. Our annotation schema is adapted from the previous English-centric works on emotion detection (Mohammad et al., 2018; Mohammad, 2022) guidelines. The dataset was created through crowdsourcing using the Toloka.ai platform ensuring high-quality of the annotation process. Then, we evaluate a range of approaches on the collected dataset, starting from linguistic-based baselines, synthetic data translated from English, to large language models (LLMs). Our findings highlight the challenges of emotion classification in non-mainstream languages like Ukrainian and emphasize the need for further development of Ukrainian-specific models and training resources.
title EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian
topic Computation and Language
url https://arxiv.org/abs/2505.23297