OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
Fuente:
arXiv
Saved in:
| Main Authors: | Kanerva, Jenna, Ledins, Cassandra, Käpyaho, Siiri, Ginter, Filip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs
by: Laato, Joonatan, et al.
Published: (2025)
by: Laato, Joonatan, et al.
Published: (2025)
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
by: Myntti, Amanda, et al.
Published: (2026)
by: Myntti, Amanda, et al.
Published: (2026)
Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs
by: Laato, Joonatan, et al.
Published: (2026)
by: Laato, Joonatan, et al.
Published: (2026)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025)
by: Greif, Gavin, et al.
Published: (2025)
Semantic Search as Extractive Paraphrase Span Detection
by: Kanerva, Jenna, et al.
Published: (2021)
by: Kanerva, Jenna, et al.
Published: (2021)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
Creating a Historical Migration Dataset from Finnish Church Records, 1800-1920
by: Vesalainen, Ari, et al.
Published: (2025)
by: Vesalainen, Ari, et al.
Published: (2025)
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering
by: Henriksson, Erik, et al.
Published: (2025)
by: Henriksson, Erik, et al.
Published: (2025)
Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
by: Nuutinen, Emil, et al.
Published: (2025)
by: Nuutinen, Emil, et al.
Published: (2025)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024)
by: Kashid, Harshvivek, et al.
Published: (2024)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024)
by: Do, Thao, et al.
Published: (2024)
Adapting LLMs for Minimal-edit Grammatical Error Correction
by: Staruch, Ryszard, et al.
Published: (2025)
by: Staruch, Ryszard, et al.
Published: (2025)
Asking LLMs to Verify First is Almost Free Lunch
by: Wu, Shiguang, et al.
Published: (2025)
by: Wu, Shiguang, et al.
Published: (2025)
Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities
by: Levchenko, Maria
Published: (2025)
by: Levchenko, Maria
Published: (2025)
Confidence-Aware Document OCR Error Detection
by: Hemmer, Arthur, et al.
Published: (2024)
by: Hemmer, Arthur, et al.
Published: (2024)
Advancing Post-OCR Correction: A Comparative Study of Synthetic Data
by: Guan, Shuhao, et al.
Published: (2024)
by: Guan, Shuhao, et al.
Published: (2024)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
by: Boros, Emanuela, et al.
Published: (2024)
by: Boros, Emanuela, et al.
Published: (2024)
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
by: Chen, Xinrui, et al.
Published: (2025)
by: Chen, Xinrui, et al.
Published: (2025)
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users
by: Hu, Mengxuan, et al.
Published: (2024)
by: Hu, Mengxuan, et al.
Published: (2024)
Improving OCR for Historical Texts of Multiple Languages
by: Westerdijk, Hylke, et al.
Published: (2025)
by: Westerdijk, Hylke, et al.
Published: (2025)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
JaPOC: Japanese Post-OCR Correction Benchmark using Vouchers
by: Fujitake, Masato
Published: (2024)
by: Fujitake, Masato
Published: (2024)
Is It a Free Lunch for Removing Outliers during Pretraining?
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
Here's a Free Lunch: Sanitizing Backdoored Models with Model Merge
by: Arora, Ansh, et al.
Published: (2024)
by: Arora, Ansh, et al.
Published: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
by: Li, Haoran, et al.
Published: (2026)
by: Li, Haoran, et al.
Published: (2026)
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
Approaches to Analysing Historical Newspapers Using LLMs
by: Dobranić, Filip, et al.
Published: (2026)
by: Dobranić, Filip, et al.
Published: (2026)
Multi-Dimensional Evaluation of LLMs for Grammatical Error Correction
by: Labib, Adnan, et al.
Published: (2026)
by: Labib, Adnan, et al.
Published: (2026)
Evaluation of LLMs on Long-tail Entity Linking in Historical Documents
by: Boscariol, Marta, et al.
Published: (2025)
by: Boscariol, Marta, et al.
Published: (2025)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
by: Wang, Chengye, et al.
Published: (2026)
by: Wang, Chengye, et al.
Published: (2026)
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
by: Manrique-Gómez, Laura, et al.
Published: (2024)
by: Manrique-Gómez, Laura, et al.
Published: (2024)
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
by: Douillard, Arthur, et al.
Published: (2025)
by: Douillard, Arthur, et al.
Published: (2025)
Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke
by: Wu, Yu, et al.
Published: (2026)
by: Wu, Yu, et al.
Published: (2026)
Enriching Historical Records: An OCR and AI-Driven Approach for Database Integration
by: Abedi, Zahra, et al.
Published: (2025)
by: Abedi, Zahra, et al.
Published: (2025)
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
by: Chang, Trenton, et al.
Published: (2025)
by: Chang, Trenton, et al.
Published: (2025)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
by: Guan, Shuhao, et al.
Published: (2025)
by: Guan, Shuhao, et al.
Published: (2025)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
by: Pang, Qi, et al.
Published: (2024)
by: Pang, Qi, et al.
Published: (2024)
Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR Data
by: Piryani, Bhawna, et al.
Published: (2025)
by: Piryani, Bhawna, et al.
Published: (2025)
Similar Items
-
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs
by: Laato, Joonatan, et al.
Published: (2025) -
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
by: Myntti, Amanda, et al.
Published: (2026) -
Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs
by: Laato, Joonatan, et al.
Published: (2026) -
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025) -
Semantic Search as Extractive Paraphrase Span Detection
by: Kanerva, Jenna, et al.
Published: (2021)