SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Smădu, Răzvan-Alexandru, Iuga, Andreea, Cercel, Dumitru-Clementin, Pop, Florin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918161119444992
author Smădu, Răzvan-Alexandru
Iuga, Andreea
Cercel, Dumitru-Clementin
Pop, Florin
author_facet Smădu, Răzvan-Alexandru
Iuga, Andreea
Cercel, Dumitru-Clementin
Pop, Florin
contents Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more granular level, allowing satirical information to be incorporated into news articles. In this paper, we introduce the first sentence-level dataset for Romanian satire detection for news articles, called SeLeRoSa. The dataset comprises 13,873 manually annotated sentences spanning various domains, including social issues, IT, science, and movies. With the rise and recent progress of large language models (LLMs) in the natural language processing literature, LLMs have demonstrated enhanced capabilities to tackle various tasks in zero-shot settings. We evaluate multiple baseline models based on LLMs in both zero-shot and fine-tuning settings, as well as baseline transformer-based models. Our findings reveal the current limitations of these models in the sentence-level satire detection task, paving the way for new research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00893
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
Smădu, Răzvan-Alexandru
Iuga, Andreea
Cercel, Dumitru-Clementin
Pop, Florin
Computation and Language
I.2.7; I.7
Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more granular level, allowing satirical information to be incorporated into news articles. In this paper, we introduce the first sentence-level dataset for Romanian satire detection for news articles, called SeLeRoSa. The dataset comprises 13,873 manually annotated sentences spanning various domains, including social issues, IT, science, and movies. With the rise and recent progress of large language models (LLMs) in the natural language processing literature, LLMs have demonstrated enhanced capabilities to tackle various tasks in zero-shot settings. We evaluate multiple baseline models based on LLMs in both zero-shot and fine-tuning settings, as well as baseline transformer-based models. Our findings reveal the current limitations of these models in the sentence-level satire detection task, paving the way for new research directions.
title SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
topic Computation and Language
I.2.7; I.7
url https://arxiv.org/abs/2509.00893