Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Schöffel, Matthias, Arias, Esteban Garces, Wiedner, Marinus, Ruppert, Paula, Li, Meimingwei, Heumann, Christian, Aßenmacher, Matthias
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911018673766400
author Schöffel, Matthias
Arias, Esteban Garces
Wiedner, Marinus
Ruppert, Paula
Li, Meimingwei
Heumann, Christian
Aßenmacher, Matthias
author_facet Schöffel, Matthias
Arias, Esteban Garces
Wiedner, Marinus
Ruppert, Paula
Li, Meimingwei
Heumann, Christian
Aßenmacher, Matthias
contents Part-of-speech (POS) tagging remains a foundational component in natural language processing pipelines, particularly critical for historical text analysis at the intersection of computational linguistics and digital humanities. Despite significant advancements in modern large language models (LLMs) for ancient languages, their application to Medieval Romance languages presents distinctive challenges stemming from diachronic linguistic evolution, spelling variations, and labeled data scarcity. This study systematically investigates the central determinants of POS tagging performance across diverse corpora of Medieval Occitan, Medieval Spanish, and Medieval French texts, spanning biblical, hagiographical, medical, and dietary domains. Through rigorous experimentation, we evaluate how fine-tuning approaches, prompt engineering, model architectures, decoding strategies, and cross-lingual transfer learning techniques affect tagging accuracy. Our results reveal both notable limitations in LLMs' ability to process historical language variations and non-standardized spelling, as well as promising specialized techniques that effectively address the unique challenges presented by low-resource historical languages.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages
Schöffel, Matthias
Arias, Esteban Garces
Wiedner, Marinus
Ruppert, Paula
Li, Meimingwei
Heumann, Christian
Aßenmacher, Matthias
Computation and Language
Machine Learning
Part-of-speech (POS) tagging remains a foundational component in natural language processing pipelines, particularly critical for historical text analysis at the intersection of computational linguistics and digital humanities. Despite significant advancements in modern large language models (LLMs) for ancient languages, their application to Medieval Romance languages presents distinctive challenges stemming from diachronic linguistic evolution, spelling variations, and labeled data scarcity. This study systematically investigates the central determinants of POS tagging performance across diverse corpora of Medieval Occitan, Medieval Spanish, and Medieval French texts, spanning biblical, hagiographical, medical, and dietary domains. Through rigorous experimentation, we evaluate how fine-tuning approaches, prompt engineering, model architectures, decoding strategies, and cross-lingual transfer learning techniques affect tagging accuracy. Our results reveal both notable limitations in LLMs' ability to process historical language variations and non-standardized spelling, as well as promising specialized techniques that effectively address the unique challenges presented by low-resource historical languages.
title Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.17715