Scrambled text: training Language Models to correct OCR errors using synthetic data
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Bourne, Jonathan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
par: Bourne, Jonathan
Publié: (2024)
par: Bourne, Jonathan
Publié: (2024)
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
par: Bourne, Jonathan
Publié: (2025)
par: Bourne, Jonathan
Publié: (2025)
CECOR: Correction-oriented synthetic data construction for factual error correction
par: Zhu, Lei, et autres
Publié: (2026)
par: Zhu, Lei, et autres
Publié: (2026)
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
par: Pieler, Michael, et autres
Publié: (2024)
par: Pieler, Michael, et autres
Publié: (2024)
The Character Error Vector: Decomposable errors for page-level OCR evaluation
par: Bourne, Jonathan, et autres
Publié: (2026)
par: Bourne, Jonathan, et autres
Publié: (2026)
Prompting open-source and commercial language models for grammatical error correction of English learner text
par: Davis, Christopher, et autres
Publié: (2024)
par: Davis, Christopher, et autres
Publié: (2024)
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
par: Steinberg, Jonathan, et autres
Publié: (2026)
par: Steinberg, Jonathan, et autres
Publié: (2026)
Typoglycemia under the Hood: Investigating Language Models' Understanding of Scrambled Words
par: Sperduti, Gianluca, et autres
Publié: (2025)
par: Sperduti, Gianluca, et autres
Publié: (2025)
Labeling Free-text Data using Language Model Ensembles
par: Qiu, Jiaxing, et autres
Publié: (2025)
par: Qiu, Jiaxing, et autres
Publié: (2025)
Private prediction for large-scale synthetic text generation
par: Amin, Kareem, et autres
Publié: (2024)
par: Amin, Kareem, et autres
Publié: (2024)
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
par: Poznanski, Jake, et autres
Publié: (2025)
par: Poznanski, Jake, et autres
Publié: (2025)
Typhoon OCR: Open Vision-Language Model For Thai Document Extraction
par: Nonesung, Surapon, et autres
Publié: (2026)
par: Nonesung, Surapon, et autres
Publié: (2026)
Spiking the training data to correct for test set contamination
par: Wei, Johnny Tian-Zheng, et autres
Publié: (2026)
par: Wei, Johnny Tian-Zheng, et autres
Publié: (2026)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
par: Wang, Chengye, et autres
Publié: (2026)
par: Wang, Chengye, et autres
Publié: (2026)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
par: Cardoso, Gabriel Pimenta de Freitas, et autres
Publié: (2026)
par: Cardoso, Gabriel Pimenta de Freitas, et autres
Publié: (2026)
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding
par: Shi, Yuling, et autres
Publié: (2026)
par: Shi, Yuling, et autres
Publié: (2026)
Self-training from Self-memory in Data-to-text Generation
par: Ta, Hoang-Thang
Publié: (2024)
par: Ta, Hoang-Thang
Publié: (2024)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
par: Lyth, Dan, et autres
Publié: (2024)
par: Lyth, Dan, et autres
Publié: (2024)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
par: Yu, Haiyang, et autres
Publié: (2025)
par: Yu, Haiyang, et autres
Publié: (2025)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
par: Hennara, Khalil, et autres
Publié: (2025)
par: Hennara, Khalil, et autres
Publié: (2025)
Improving OCR for Historical Texts of Multiple Languages
par: Westerdijk, Hylke, et autres
Publié: (2025)
par: Westerdijk, Hylke, et autres
Publié: (2025)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
par: Banerjee, Sourav, et autres
Publié: (2024)
par: Banerjee, Sourav, et autres
Publié: (2024)
Chain-of-Though (CoT) prompting strategies for medical error detection and correction
par: Wu, Zhaolong, et autres
Publié: (2024)
par: Wu, Zhaolong, et autres
Publié: (2024)
Tag and correct: high precision post-editing approach to correction of speech recognition errors
par: Ziętkiewicz, Tomasz
Publié: (2024)
par: Ziętkiewicz, Tomasz
Publié: (2024)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
par: Kargaran, Amir Hossein, et autres
Publié: (2026)
par: Kargaran, Amir Hossein, et autres
Publié: (2026)
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
par: Ascione, Grazia Sveva, et autres
Publié: (2025)
par: Ascione, Grazia Sveva, et autres
Publié: (2025)
Self-correction is Not An Innate Capability in Language Models
par: Liu, Guangliang, et autres
Publié: (2024)
par: Liu, Guangliang, et autres
Publié: (2024)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
par: He, Jie, et autres
Publié: (2025)
par: He, Jie, et autres
Publié: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
par: Poznanski, Jake, et autres
Publié: (2025)
par: Poznanski, Jake, et autres
Publié: (2025)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
par: Gagnier, Henry, et autres
Publié: (2026)
par: Gagnier, Henry, et autres
Publié: (2026)
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency
par: Liang, Mingliang, et autres
Publié: (2024)
par: Liang, Mingliang, et autres
Publié: (2024)
GLM-OCR Technical Report
par: Duan, Shuaiqi, et autres
Publié: (2026)
par: Duan, Shuaiqi, et autres
Publié: (2026)
Improving the quality of Persian clinical text with a novel spelling correction system
par: Dashti, Seyed Mohammad Sadegh, et autres
Publié: (2024)
par: Dashti, Seyed Mohammad Sadegh, et autres
Publié: (2024)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
par: Kashid, Harshvivek, et autres
Publié: (2024)
par: Kashid, Harshvivek, et autres
Publié: (2024)
Investigating the translation capabilities of Large Language Models trained on parallel data only
par: Gilabert, Javier García, et autres
Publié: (2024)
par: Gilabert, Javier García, et autres
Publié: (2024)
Audio-visual training for improved grounding in video-text LLMs
par: Sagare, Shivprasad, et autres
Publié: (2024)
par: Sagare, Shivprasad, et autres
Publié: (2024)
Spanish TrOCR: Leveraging Transfer Learning for Language Adaptation
par: Lauar, Filipe, et autres
Publié: (2024)
par: Lauar, Filipe, et autres
Publié: (2024)
Causality extraction from medical text using Large Language Models (LLMs)
par: Gopalakrishnan, Seethalakshmi, et autres
Publié: (2024)
par: Gopalakrishnan, Seethalakshmi, et autres
Publié: (2024)
Jochre 3 and the Yiddish OCR corpus
par: Urieli, Assaf, et autres
Publié: (2025)
par: Urieli, Assaf, et autres
Publié: (2025)
Transferable text data distillation by trajectory matching
par: Yao, Rong, et autres
Publié: (2025)
par: Yao, Rong, et autres
Publié: (2025)
Documents similaires
-
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
par: Bourne, Jonathan
Publié: (2024) -
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
par: Bourne, Jonathan
Publié: (2025) -
CECOR: Correction-oriented synthetic data construction for factual error correction
par: Zhu, Lei, et autres
Publié: (2026) -
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
par: Pieler, Michael, et autres
Publié: (2024) -
The Character Error Vector: Decomposable errors for page-level OCR evaluation
par: Bourne, Jonathan, et autres
Publié: (2026)