Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Greif, Gavin, Griesshaber, Niclas, Greif, Robin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards the AI Historian: Agentic Information Extraction from Primary Sources
by: Hufe, Lorenz, et al.
Published: (2026)
by: Hufe, Lorenz, et al.
Published: (2026)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
Named Entity Recognition of Historical Texts via Large Language Model
by: Zhang, Shibingfeng, et al.
Published: (2025)
by: Zhang, Shibingfeng, et al.
Published: (2025)
Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
by: Griesshaber, Niclas, et al.
Published: (2025)
by: Griesshaber, Niclas, et al.
Published: (2025)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
by: Boros, Emanuela, et al.
Published: (2024)
by: Boros, Emanuela, et al.
Published: (2024)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)
by: Karamolegkou, Antonia, et al.
Published: (2026)
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
by: Bourne, Jonathan
Published: (2024)
by: Bourne, Jonathan
Published: (2024)
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
by: Manrique-Gómez, Laura, et al.
Published: (2024)
by: Manrique-Gómez, Laura, et al.
Published: (2024)
Layout-Aware OCR for Black Digital Archives with Unsupervised Evaluation
by: Beyene, Fitsum Sileshi, et al.
Published: (2025)
by: Beyene, Fitsum Sileshi, et al.
Published: (2025)
Hybrid X-Linker: Automated Data Generation and Extreme Multi-label Ranking for Biomedical Entity Linking
by: Ruas, Pedro, et al.
Published: (2024)
by: Ruas, Pedro, et al.
Published: (2024)
PubMed-OCR: PMC Open Access OCR Annotations
by: Heidenreich, Hunter, et al.
Published: (2026)
by: Heidenreich, Hunter, et al.
Published: (2026)
A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
by: Beyene, Fitsum Sileshi, et al.
Published: (2026)
by: Beyene, Fitsum Sileshi, et al.
Published: (2026)
Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
LLMs4Synthesis: Leveraging Large Language Models for Scientific Synthesis
by: Giglou, Hamed Babaei, et al.
Published: (2024)
by: Giglou, Hamed Babaei, et al.
Published: (2024)
Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents
by: Arcan, Mihael
Published: (2025)
by: Arcan, Mihael
Published: (2025)
Accelerating Scientific Discovery with Multi-Document Summarization of Impact-Ranked Papers
by: Koloveas, Paris, et al.
Published: (2025)
by: Koloveas, Paris, et al.
Published: (2025)
LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models
by: Sadruddin, Sameer, et al.
Published: (2025)
by: Sadruddin, Sameer, et al.
Published: (2025)
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
by: McCoubrey, Michael, et al.
Published: (2026)
by: McCoubrey, Michael, et al.
Published: (2026)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
by: Seo, Yongsik, et al.
Published: (2026)
by: Seo, Yongsik, et al.
Published: (2026)
Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish
by: Cohen, Kevin, et al.
Published: (2025)
by: Cohen, Kevin, et al.
Published: (2025)
Large language models for automated scholarly paper review: A survey
by: Zhuang, Zhenzhen, et al.
Published: (2025)
by: Zhuang, Zhenzhen, et al.
Published: (2025)
The Attribution Crisis in LLM Search Results
by: Strauss, Ilan, et al.
Published: (2025)
by: Strauss, Ilan, et al.
Published: (2025)
Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
SciClaims: An End-to-End Generative System for Biomedical Claim Analysis
by: Ortega, Raúl, et al.
Published: (2025)
by: Ortega, Raúl, et al.
Published: (2025)
SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
by: Wu, Wenqing, et al.
Published: (2025)
by: Wu, Wenqing, et al.
Published: (2025)
Fairness Evaluation of Large Language Models in Academic Library Reference Services
by: Wang, Haining, et al.
Published: (2025)
by: Wang, Haining, et al.
Published: (2025)
Polarity Detection of Sustainable Development Goals in News Text
by: Cadeddu, Andrea, et al.
Published: (2025)
by: Cadeddu, Andrea, et al.
Published: (2025)
Annotating Scientific Uncertainty: A comprehensive model using linguistic patterns and comparison with existing approaches
by: Ningrum, Panggih Kusuma, et al.
Published: (2025)
by: Ningrum, Panggih Kusuma, et al.
Published: (2025)
Big Tech-Funded AI Papers Have Higher Citation Impact, Greater Insularity, and Larger Recency Bias
by: Gnewuch, Max Martin, et al.
Published: (2025)
by: Gnewuch, Max Martin, et al.
Published: (2025)
Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models
by: Tan, Zhiyin, et al.
Published: (2025)
by: Tan, Zhiyin, et al.
Published: (2025)
Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
by: Tan, Zhiyin, et al.
Published: (2025)
by: Tan, Zhiyin, et al.
Published: (2025)
Context Selection for Hypothesis and Statistical Evidence Extraction from Full-Text Scientific Articles
by: Koneru, Sai, et al.
Published: (2026)
by: Koneru, Sai, et al.
Published: (2026)
SPACE-IDEAS: A Dataset for Salient Information Detection in Space Innovation
by: García-Silva, Andrés, et al.
Published: (2024)
by: García-Silva, Andrés, et al.
Published: (2024)
Scientific Statement Classification over arXiv.org
by: Ginev, Deyan, et al.
Published: (2019)
by: Ginev, Deyan, et al.
Published: (2019)
NewsEdits 2.0: Learning the Intentions Behind Updating News
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
Textual Entailment for Effective Triple Validation in Object Prediction
by: García-Silva, Andrés, et al.
Published: (2024)
by: García-Silva, Andrés, et al.
Published: (2024)
NAIST Academic Travelogue Dataset
by: Ouchi, Hiroki, et al.
Published: (2023)
by: Ouchi, Hiroki, et al.
Published: (2023)
Exploring Scholarly Data by Semantic Query on Knowledge Graph Embedding Space
by: Tran, Hung Nghiep, et al.
Published: (2019)
by: Tran, Hung Nghiep, et al.
Published: (2019)
Similar Items
-
Towards the AI Historian: Agentic Information Extraction from Primary Sources
by: Hufe, Lorenz, et al.
Published: (2026) -
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024) -
Named Entity Recognition of Historical Texts via Large Language Model
by: Zhang, Shibingfeng, et al.
Published: (2025) -
Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
by: Griesshaber, Niclas, et al.
Published: (2025) -
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
by: Boros, Emanuela, et al.
Published: (2024)