Improving OCR for Historical Texts of Multiple Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Westerdijk, Hylke, Blankenborg, Ben, Islam, Khondoker Ittehadul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916898429468672
author Westerdijk, Hylke
Blankenborg, Ben
Islam, Khondoker Ittehadul
author_facet Westerdijk, Hylke
Blankenborg, Ben
Islam, Khondoker Ittehadul
contents This paper presents our methodology and findings from three tasks across Optical Character Recognition (OCR) and Document Layout Analysis using advanced deep learning techniques. First, for the historical Hebrew fragments of the Dead Sea Scrolls, we enhanced our dataset through extensive data augmentation and employed the Kraken and TrOCR models to improve character recognition. In our analysis of 16th to 18th-century meeting resolutions task, we utilized a Convolutional Recurrent Neural Network (CRNN) that integrated DeepLabV3+ for semantic segmentation with a Bidirectional LSTM, incorporating confidence-based pseudolabeling to refine our model. Finally, for modern English handwriting recognition task, we applied a CRNN with a ResNet34 encoder, trained using the Connectionist Temporal Classification (CTC) loss function to effectively capture sequential dependencies. This report offers valuable insights and suggests potential directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving OCR for Historical Texts of Multiple Languages
Westerdijk, Hylke
Blankenborg, Ben
Islam, Khondoker Ittehadul
Computer Vision and Pattern Recognition
Computation and Language
This paper presents our methodology and findings from three tasks across Optical Character Recognition (OCR) and Document Layout Analysis using advanced deep learning techniques. First, for the historical Hebrew fragments of the Dead Sea Scrolls, we enhanced our dataset through extensive data augmentation and employed the Kraken and TrOCR models to improve character recognition. In our analysis of 16th to 18th-century meeting resolutions task, we utilized a Convolutional Recurrent Neural Network (CRNN) that integrated DeepLabV3+ for semantic segmentation with a Bidirectional LSTM, incorporating confidence-based pseudolabeling to refine our model. Finally, for modern English handwriting recognition task, we applied a CRNN with a ResNet34 encoder, trained using the Connectionist Temporal Classification (CTC) loss function to effectively capture sequential dependencies. This report offers valuable insights and suggests potential directions for future research.
title Improving OCR for Historical Texts of Multiple Languages
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2508.10356