Advancing Post-OCR Correction: A Comparative Study of Synthetic Data
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, Shuhao, Greene, Derek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
by: Guan, Shuhao, et al.
Published: (2025)
by: Guan, Shuhao, et al.
Published: (2025)
Benchmark Data Contamination of Large Language Models: A Survey
by: Xu, Cheng, et al.
Published: (2024)
by: Xu, Cheng, et al.
Published: (2024)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024)
by: Kashid, Harshvivek, et al.
Published: (2024)
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
by: Kanerva, Jenna, et al.
Published: (2025)
by: Kanerva, Jenna, et al.
Published: (2025)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025)
by: Greif, Gavin, et al.
Published: (2025)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
JaPOC: Japanese Post-OCR Correction Benchmark using Vouchers
by: Fujitake, Masato
Published: (2024)
by: Fujitake, Masato
Published: (2024)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
by: Gagnier, Henry, et al.
Published: (2026)
by: Gagnier, Henry, et al.
Published: (2026)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
by: Wang, Chengye, et al.
Published: (2026)
by: Wang, Chengye, et al.
Published: (2026)
Advances and Limitations in Open Source Arabic-Script OCR: A Case Study
by: Kiessling, Benjamin, et al.
Published: (2024)
by: Kiessling, Benjamin, et al.
Published: (2024)
Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian
by: Samo, Giuseppe, et al.
Published: (2026)
by: Samo, Giuseppe, et al.
Published: (2026)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil
by: Jayatilleke, Nevidu, et al.
Published: (2025)
by: Jayatilleke, Nevidu, et al.
Published: (2025)
DCR: Quantifying Data Contamination in LLMs Evaluation
by: Xu, Cheng, et al.
Published: (2025)
by: Xu, Cheng, et al.
Published: (2025)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
by: Cunningham, Eoghan, et al.
Published: (2025)
by: Cunningham, Eoghan, et al.
Published: (2025)
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
by: Bourne, Jonathan
Published: (2024)
by: Bourne, Jonathan
Published: (2024)
Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR Data
by: Piryani, Bhawna, et al.
Published: (2025)
by: Piryani, Bhawna, et al.
Published: (2025)
Synthetic Clarification and Correction Dialogues about Data-Centric Tasks -- A Teacher-Student Approach
by: Poelitz, Christian, et al.
Published: (2025)
by: Poelitz, Christian, et al.
Published: (2025)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
by: Al-Homoud, Haneen, et al.
Published: (2025)
by: Al-Homoud, Haneen, et al.
Published: (2025)
Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation
by: Cunningham, Eoghan, et al.
Published: (2026)
by: Cunningham, Eoghan, et al.
Published: (2026)
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
Advancing Annotation of Stance in Social Media Posts: A Comparative Analysis of Large Language Models and Crowd Sourcing
by: Li, Mao, et al.
Published: (2024)
by: Li, Mao, et al.
Published: (2024)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024)
by: Do, Thao, et al.
Published: (2024)
GLM-OCR Technical Report
by: Duan, Shuaiqi, et al.
Published: (2026)
by: Duan, Shuaiqi, et al.
Published: (2026)
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
Jochre 3 and the Yiddish OCR corpus
by: Urieli, Assaf, et al.
Published: (2025)
by: Urieli, Assaf, et al.
Published: (2025)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
by: Gill, Waris, et al.
Published: (2025)
by: Gill, Waris, et al.
Published: (2025)
Towards the Development of Balanced Synthetic Data for Correcting Grammatical Errors in Arabic: An Approach Based on Error Tagging Model and Synthetic Data Generating Model
by: Alrehili, Ahlam, et al.
Published: (2025)
by: Alrehili, Ahlam, et al.
Published: (2025)
JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching
by: Magron, Antoine, et al.
Published: (2024)
by: Magron, Antoine, et al.
Published: (2024)
Cleansing Jewel: A Neural Spelling Correction Model Built On Google OCR-ed Tibetan Manuscripts
by: Luo, Queenie, et al.
Published: (2023)
by: Luo, Queenie, et al.
Published: (2023)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts
by: Bagdon, Christopher, et al.
Published: (2025)
by: Bagdon, Christopher, et al.
Published: (2025)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
by: Kargaran, Amir Hossein, et al.
Published: (2026)
by: Kargaran, Amir Hossein, et al.
Published: (2026)
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning
by: Fang, Tao, et al.
Published: (2024)
by: Fang, Tao, et al.
Published: (2024)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025)
by: Zhang, Yizhuo, et al.
Published: (2025)
An Empirical Study of Validating Synthetic Data for Formula Generation
by: Singh, Usneek, et al.
Published: (2024)
by: Singh, Usneek, et al.
Published: (2024)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
by: Yılmaz, Deniz, et al.
Published: (2026)
by: Yılmaz, Deniz, et al.
Published: (2026)
A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic
by: Gonzalez, Juan Moreno, et al.
Published: (2025)
by: Gonzalez, Juan Moreno, et al.
Published: (2025)
Advancing Semi-Supervised Learning for Automatic Post-Editing: Data-Synthesis by Mask-Infilling with Erroneous Terms
by: Lee, Wonkee, et al.
Published: (2022)
by: Lee, Wonkee, et al.
Published: (2022)
Similar Items
-
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
by: Guan, Shuhao, et al.
Published: (2025) -
Benchmark Data Contamination of Large Language Models: A Survey
by: Xu, Cheng, et al.
Published: (2024) -
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024) -
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
by: Kanerva, Jenna, et al.
Published: (2025) -
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025)