Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Kevin, Manrique-Gómez, Laura, Manrique, Rubén
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915217915510784
author Cohen, Kevin
Manrique-Gómez, Laura
Manrique, Rubén
author_facet Cohen, Kevin
Manrique-Gómez, Laura
Manrique, Rubén
contents This study explores the use of large language models (LLMs) to enhance datasets and improve irony detection in 19th-century Latin American newspapers. Two strategies were employed to evaluate the efficacy of BERT and GPT-4o models in capturing the subtle nuances nature of irony, through both multi-class and binary classification tasks. First, we implemented dataset enhancements focused on enriching emotional and contextual cues; however, these showed limited impact on historical language analysis. The second strategy, a semi-automated annotation process, effectively addressed class imbalance and augmented the dataset with high-quality annotations. Despite the challenges posed by the complexity of irony, this work contributes to the advancement of sentiment analysis through two key contributions: introducing a new historical Spanish dataset tagged for sentiment analysis and irony detection, and proposing a semi-automated annotation methodology where human expertise is crucial for refining LLMs results, enriched by incorporating historical and cultural contexts as core features.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish
Cohen, Kevin
Manrique-Gómez, Laura
Manrique, Rubén
Computation and Language
Artificial Intelligence
Digital Libraries
I.2.7
This study explores the use of large language models (LLMs) to enhance datasets and improve irony detection in 19th-century Latin American newspapers. Two strategies were employed to evaluate the efficacy of BERT and GPT-4o models in capturing the subtle nuances nature of irony, through both multi-class and binary classification tasks. First, we implemented dataset enhancements focused on enriching emotional and contextual cues; however, these showed limited impact on historical language analysis. The second strategy, a semi-automated annotation process, effectively addressed class imbalance and augmented the dataset with high-quality annotations. Despite the challenges posed by the complexity of irony, this work contributes to the advancement of sentiment analysis through two key contributions: introducing a new historical Spanish dataset tagged for sentiment analysis and irony detection, and proposing a semi-automated annotation methodology where human expertise is crucial for refining LLMs results, enriched by incorporating historical and cultural contexts as core features.
title Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish
topic Computation and Language
Artificial Intelligence
Digital Libraries
I.2.7
url https://arxiv.org/abs/2503.22585