Improving Automatic Text Recognition with Language Models in the PyLaia Open-Source Library

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tarride, Solène, Schneider, Yoann, Generali-Lince, Marie, Boillet, Mélodie, Abadie, Bastien, Kermorvant, Christopher
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917652685914112
author Tarride, Solène
Schneider, Yoann
Generali-Lince, Marie
Boillet, Mélodie
Abadie, Bastien
Kermorvant, Christopher
author_facet Tarride, Solène
Schneider, Yoann
Generali-Lince, Marie
Boillet, Mélodie
Abadie, Bastien
Kermorvant, Christopher
contents PyLaia is one of the most popular open-source software for Automatic Text Recognition (ATR), delivering strong performance in terms of speed and accuracy. In this paper, we outline our recent contributions to the PyLaia library, focusing on the incorporation of reliable confidence scores and the integration of statistical language modeling during decoding. Our implementation provides an easy way to combine PyLaia with n-grams language models at different levels. One of the highlights of this work is that language models are completely auto-tuned: they can be built and used easily without any expert knowledge, and without requiring any additional data. To demonstrate the significance of our contribution, we evaluate PyLaia's performance on twelve datasets, both with and without language modelling. The results show that decoding with small language models improves the Word Error Rate by 13% and the Character Error Rate by 12% in average. Additionally, we conduct an analysis of confidence scores and highlight the importance of calibration techniques. Our implementation is publicly available in the official PyLaia repository at https://gitlab.teklia.com/atr/pylaia, and twelve open-source models are released on Hugging Face.
format Preprint
id arxiv_https___arxiv_org_abs_2404_18722
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Automatic Text Recognition with Language Models in the PyLaia Open-Source Library
Tarride, Solène
Schneider, Yoann
Generali-Lince, Marie
Boillet, Mélodie
Abadie, Bastien
Kermorvant, Christopher
Computer Vision and Pattern Recognition
Computation and Language
PyLaia is one of the most popular open-source software for Automatic Text Recognition (ATR), delivering strong performance in terms of speed and accuracy. In this paper, we outline our recent contributions to the PyLaia library, focusing on the incorporation of reliable confidence scores and the integration of statistical language modeling during decoding. Our implementation provides an easy way to combine PyLaia with n-grams language models at different levels. One of the highlights of this work is that language models are completely auto-tuned: they can be built and used easily without any expert knowledge, and without requiring any additional data. To demonstrate the significance of our contribution, we evaluate PyLaia's performance on twelve datasets, both with and without language modelling. The results show that decoding with small language models improves the Word Error Rate by 13% and the Character Error Rate by 12% in average. Additionally, we conduct an analysis of confidence scores and highlight the importance of calibration techniques. Our implementation is publicly available in the official PyLaia repository at https://gitlab.teklia.com/atr/pylaia, and twelve open-source models are released on Hugging Face.
title Improving Automatic Text Recognition with Language Models in the PyLaia Open-Source Library
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2404.18722