Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915217915510784 |
|---|---|
| author | Cohen, Kevin Manrique-Gómez, Laura Manrique, Rubén |
| author_facet | Cohen, Kevin Manrique-Gómez, Laura Manrique, Rubén |
| contents | This study explores the use of large language models (LLMs) to enhance datasets and improve irony detection in 19th-century Latin American newspapers. Two strategies were employed to evaluate the efficacy of BERT and GPT-4o models in capturing the subtle nuances nature of irony, through both multi-class and binary classification tasks. First, we implemented dataset enhancements focused on enriching emotional and contextual cues; however, these showed limited impact on historical language analysis. The second strategy, a semi-automated annotation process, effectively addressed class imbalance and augmented the dataset with high-quality annotations. Despite the challenges posed by the complexity of irony, this work contributes to the advancement of sentiment analysis through two key contributions: introducing a new historical Spanish dataset tagged for sentiment analysis and irony detection, and proposing a semi-automated annotation methodology where human expertise is crucial for refining LLMs results, enriched by incorporating historical and cultural contexts as core features. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_22585 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish Cohen, Kevin Manrique-Gómez, Laura Manrique, Rubén Computation and Language Artificial Intelligence Digital Libraries I.2.7 This study explores the use of large language models (LLMs) to enhance datasets and improve irony detection in 19th-century Latin American newspapers. Two strategies were employed to evaluate the efficacy of BERT and GPT-4o models in capturing the subtle nuances nature of irony, through both multi-class and binary classification tasks. First, we implemented dataset enhancements focused on enriching emotional and contextual cues; however, these showed limited impact on historical language analysis. The second strategy, a semi-automated annotation process, effectively addressed class imbalance and augmented the dataset with high-quality annotations. Despite the challenges posed by the complexity of irony, this work contributes to the advancement of sentiment analysis through two key contributions: introducing a new historical Spanish dataset tagged for sentiment analysis and irony detection, and proposing a semi-automated annotation methodology where human expertise is crucial for refining LLMs results, enriched by incorporating historical and cultural contexts as core features. |
| title | Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish |
| topic | Computation and Language Artificial Intelligence Digital Libraries I.2.7 |
| url | https://arxiv.org/abs/2503.22585 |