Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yu, Shu, Ke, Fischer, Jonas, Pivovarova, Lidia, Rosson, David, Mäkelä, Eetu, Tolonen, Mikko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
by: Vesalainen, Ari, et al.
Published: (2026)
by: Vesalainen, Ari, et al.
Published: (2026)
Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke
by: Wu, Yu, et al.
Published: (2026)
by: Wu, Yu, et al.
Published: (2026)
Named Entity Recognition of Historical Texts via Large Language Model
by: Zhang, Shibingfeng, et al.
Published: (2025)
by: Zhang, Shibingfeng, et al.
Published: (2025)
Hidden Entity Detection from GitHub Leveraging Large Language Models
by: Gan, Lu, et al.
Published: (2025)
by: Gan, Lu, et al.
Published: (2025)
Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature
by: Schelb, Julian, et al.
Published: (2026)
by: Schelb, Julian, et al.
Published: (2026)
VisTaxa: Developing a Taxonomy of Historical Visualizations
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
ZuantuSet: A Collection of Historical Chinese Visualizations and Illustrations
by: Mei, Xiyao, et al.
Published: (2025)
by: Mei, Xiyao, et al.
Published: (2025)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025)
by: Greif, Gavin, et al.
Published: (2025)
OldVisOnline: Curating a Dataset of Historical Visualizations
by: Zhang, Yu, et al.
Published: (2023)
by: Zhang, Yu, et al.
Published: (2023)
Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish
by: Cohen, Kevin, et al.
Published: (2025)
by: Cohen, Kevin, et al.
Published: (2025)
A Stylometric Application of Large Language Models
by: Stropkay, Harrison F., et al.
Published: (2025)
by: Stropkay, Harrison F., et al.
Published: (2025)
Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents
by: Humphries, Mark, et al.
Published: (2024)
by: Humphries, Mark, et al.
Published: (2024)
PST-Bench: Tracing and Benchmarking the Source of Publications
by: Zhang, Fanjin, et al.
Published: (2024)
by: Zhang, Fanjin, et al.
Published: (2024)
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
by: Manrique-Gómez, Laura, et al.
Published: (2024)
by: Manrique-Gómez, Laura, et al.
Published: (2024)
Internal and External Impacts of Natural Language Processing Papers
by: Zhang, Yu
Published: (2025)
by: Zhang, Yu
Published: (2025)
AutoLLM-CARD: Towards a Description and Landscape of Large Language Models
by: Tian, Shengwei, et al.
Published: (2024)
by: Tian, Shengwei, et al.
Published: (2024)
LLM4SR: A Survey on Large Language Models for Scientific Research
by: Luo, Ziming, et al.
Published: (2025)
by: Luo, Ziming, et al.
Published: (2025)
"Don't Teach Minerva": Guiding LLMs Through Complex Syntax for Faithful Latin Translation with RAG
by: Aguilar, Sergio Torres
Published: (2025)
by: Aguilar, Sergio Torres
Published: (2025)
LyCon: Lyrics Reconstruction from the Bag-of-Words Using Large Language Models
by: Kim, Haven, et al.
Published: (2024)
by: Kim, Haven, et al.
Published: (2024)
Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
by: Aggarwal, Tanay, et al.
Published: (2025)
by: Aggarwal, Tanay, et al.
Published: (2025)
More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research
by: Buskirk, Trent D., et al.
Published: (2025)
by: Buskirk, Trent D., et al.
Published: (2025)
Rdgai: Classifying transcriptional changes using Large Language Models with a test case from an Arabic Gospel tradition
by: Turnbull, Robert
Published: (2025)
by: Turnbull, Robert
Published: (2025)
Relying on recent and temporally dispersed science predicts breakthrough inventions
by: Ke, Qing, et al.
Published: (2021)
by: Ke, Qing, et al.
Published: (2021)
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
by: Cargnelutti, Matteo, et al.
Published: (2025)
by: Cargnelutti, Matteo, et al.
Published: (2025)
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
by: Qin, Chuan, et al.
Published: (2025)
by: Qin, Chuan, et al.
Published: (2025)
EDDA-Coordinata: An Annotated Dataset of Historical Geographic Coordinates
by: Moncla, Ludovic, et al.
Published: (2026)
by: Moncla, Ludovic, et al.
Published: (2026)
Structured Analysis and Comparison of Alphabets in Historical Handwritten Ciphers
by: Méndez, Martín, et al.
Published: (2024)
by: Méndez, Martín, et al.
Published: (2024)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023)
by: Movva, Rajiv, et al.
Published: (2023)
Towards understanding evolution of science through language model series
by: Dong, Junjie, et al.
Published: (2024)
by: Dong, Junjie, et al.
Published: (2024)
Language Models Should be Used to Surface the Unwritten Code of Science and Society
by: Bao, Honglin, et al.
Published: (2025)
by: Bao, Honglin, et al.
Published: (2025)
A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
by: Beyene, Fitsum Sileshi, et al.
Published: (2026)
by: Beyene, Fitsum Sileshi, et al.
Published: (2026)
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
by: Shi, Kaiwen, et al.
Published: (2026)
by: Shi, Kaiwen, et al.
Published: (2026)
Automatic Detection of Research Values from Scientific Abstracts Across Computer Science Subfields
by: Jiang, Hang, et al.
Published: (2025)
by: Jiang, Hang, et al.
Published: (2025)
Trends in Equal-Contribution Authorship: A Large-Scale Bibliometric Analysis of Biomedical Literature
by: Xu, Binbin
Published: (2026)
by: Xu, Binbin
Published: (2026)
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts
by: Shahi, Gautam Kishore, et al.
Published: (2025)
by: Shahi, Gautam Kishore, et al.
Published: (2025)
Identity resolution of software metadata using Large Language Models
by: del Pico, Eva Martín, et al.
Published: (2025)
by: del Pico, Eva Martín, et al.
Published: (2025)
Using General Large Language Models to Classify Mathematical Documents
by: Ion, Patrick D. F., et al.
Published: (2024)
by: Ion, Patrick D. F., et al.
Published: (2024)
From Division to Unity: A Large-Scale Study on the Emergence of Computational Social Science, 1990-2021
by: Bao, Honglin, et al.
Published: (2024)
by: Bao, Honglin, et al.
Published: (2024)
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
by: Bourne, Jonathan
Published: (2024)
by: Bourne, Jonathan
Published: (2024)
Similar Items
-
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
by: Vesalainen, Ari, et al.
Published: (2026) -
Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke
by: Wu, Yu, et al.
Published: (2026) -
Named Entity Recognition of Historical Texts via Large Language Model
by: Zhang, Shibingfeng, et al.
Published: (2025) -
Hidden Entity Detection from GitHub Leveraging Large Language Models
by: Gan, Lu, et al.
Published: (2025) -
Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature
by: Schelb, Julian, et al.
Published: (2026)