Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908817040605184 |
|---|---|
| author | Wu, Yu Shu, Ke Fischer, Jonas Pivovarova, Lidia Rosson, David Mäkelä, Eetu Tolonen, Mikko |
| author_facet | Wu, Yu Shu, Ke Fischer, Jonas Pivovarova, Lidia Rosson, David Mäkelä, Eetu Tolonen, Mikko |
| contents | This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_19585 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark Wu, Yu Shu, Ke Fischer, Jonas Pivovarova, Lidia Rosson, David Mäkelä, Eetu Tolonen, Mikko Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Digital Libraries This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin. |
| title | Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Digital Libraries |
| url | https://arxiv.org/abs/2510.19585 |