Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Santini, Cristian, Melosi, Laura, Frontoni, Emanuele
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908380247883776
author Santini, Cristian
Melosi, Laura
Frontoni, Emanuele
author_facet Santini, Cristian
Melosi, Laura
Frontoni, Emanuele
contents The increased digitization of world's textual heritage poses significant challenges for both computer science and literary studies. Overall, there is an urgent need of computational techniques able to adapt to the challenges of historical texts, such as orthographic and spelling variations, fragmentary structure and digitization errors. The rise of large language models (LLMs) has revolutionized natural language processing, suggesting promising applications for Named Entity Recognition (NER) on historical documents. In spite of this, no thorough evaluation has been proposed for Italian texts. This research tries to fill the gap by proposing a new challenging dataset for entity extraction based on a corpus of 19th century scholarly notes, i.e. Giacomo Leopardi's Zibaldone (1898), containing 2,899 references to people, locations and literary works. This dataset was used to carry out reproducible experiments with both domain-specific BERT-based models and state-of-the-art LLMs such as LLaMa3.1. Results show that instruction-tuned models encounter multiple difficulties handling historical humanistic texts, while fine-tuned NER models offer more robust performance even with challenging entity types such as bibliographic references.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone
Santini, Cristian
Melosi, Laura
Frontoni, Emanuele
Computation and Language
Artificial Intelligence
The increased digitization of world's textual heritage poses significant challenges for both computer science and literary studies. Overall, there is an urgent need of computational techniques able to adapt to the challenges of historical texts, such as orthographic and spelling variations, fragmentary structure and digitization errors. The rise of large language models (LLMs) has revolutionized natural language processing, suggesting promising applications for Named Entity Recognition (NER) on historical documents. In spite of this, no thorough evaluation has been proposed for Italian texts. This research tries to fill the gap by proposing a new challenging dataset for entity extraction based on a corpus of 19th century scholarly notes, i.e. Giacomo Leopardi's Zibaldone (1898), containing 2,899 references to people, locations and literary works. This dataset was used to carry out reproducible experiments with both domain-specific BERT-based models and state-of-the-art LLMs such as LLaMa3.1. Results show that instruction-tuned models encounter multiple difficulties handling historical humanistic texts, while fine-tuned NER models offer more robust performance even with challenging entity types such as bibliographic references.
title Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.20113