Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Yu, Shu, Ke, Fischer, Jonas, Pivovarova, Lidia, Rosson, David, Mäkelä, Eetu, Tolonen, Mikko
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908817040605184
author Wu, Yu
Shu, Ke
Fischer, Jonas
Pivovarova, Lidia
Rosson, David
Mäkelä, Eetu
Tolonen, Mikko
author_facet Wu, Yu
Shu, Ke
Fischer, Jonas
Pivovarova, Lidia
Rosson, David
Mäkelä, Eetu
Tolonen, Mikko
contents This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
Wu, Yu
Shu, Ke
Fischer, Jonas
Pivovarova, Lidia
Rosson, David
Mäkelä, Eetu
Tolonen, Mikko
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Digital Libraries
This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin.
title Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Digital Libraries
url https://arxiv.org/abs/2510.19585