Understanding Cross-Language Transfer Improvements in Low-Resource HTR: The Role of Sequence Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Al-azzawi, Sana, Liu, Chang, Habib, Nudrat, Barney, Elisa, Liwicki, Marcus
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918487133257728
author Al-azzawi, Sana
Liu, Chang
Habib, Nudrat
Barney, Elisa
Liwicki, Marcus
author_facet Al-azzawi, Sana
Liu, Chang
Habib, Nudrat
Barney, Elisa
Liwicki, Marcus
contents Handwritten Text Recognition (HTR) for Arabic-script languages benefits from cross-language joint training under low-resource conditions, particularly when using CRNN-based models that combine convolutional encoders with sequence modeling. However, it remains unclear whether these improvements are better explained by shared visual representations or sequence-level dependencies. In this work, we conduct a controlled architectural study of line-level Arabic-script HTR, comparing CNN-only models with CTC decoding and CRNN models under identical single-script and multi-script training regimes. Experiments are performed on Arabic (KHATT), Urdu (NUST-UHWR), and Persian (PHTD) datasets under low-resource settings (K in {100, 500, 1000}). Our results show a clear divergence in transfer behavior: while CNN-only models exhibit limited or unstable improvements, CRNN models achieve better performance under multi-script training, particularly in the most data-constrained regimes. Focusing on transfer improvements (delta CER) rather than absolute performance, we find that cross-language improvements are associated with sequence-level modeling, while sharing visual representations learned by the CNN encoder, corresponding to similarities in character shapes across scripts, alone appears to be insufficient. This finding suggests that contextual modeling plays an important role in enabling effective transfer in low-resource scenarios, and that similar behavior may extend to other low-resource language settings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05900
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Cross-Language Transfer Improvements in Low-Resource HTR: The Role of Sequence Modeling
Al-azzawi, Sana
Liu, Chang
Habib, Nudrat
Barney, Elisa
Liwicki, Marcus
Computer Vision and Pattern Recognition
Handwritten Text Recognition (HTR) for Arabic-script languages benefits from cross-language joint training under low-resource conditions, particularly when using CRNN-based models that combine convolutional encoders with sequence modeling. However, it remains unclear whether these improvements are better explained by shared visual representations or sequence-level dependencies. In this work, we conduct a controlled architectural study of line-level Arabic-script HTR, comparing CNN-only models with CTC decoding and CRNN models under identical single-script and multi-script training regimes. Experiments are performed on Arabic (KHATT), Urdu (NUST-UHWR), and Persian (PHTD) datasets under low-resource settings (K in {100, 500, 1000}). Our results show a clear divergence in transfer behavior: while CNN-only models exhibit limited or unstable improvements, CRNN models achieve better performance under multi-script training, particularly in the most data-constrained regimes. Focusing on transfer improvements (delta CER) rather than absolute performance, we find that cross-language improvements are associated with sequence-level modeling, while sharing visual representations learned by the CNN encoder, corresponding to similarities in character shapes across scripts, alone appears to be insufficient. This finding suggests that contextual modeling plays an important role in enabling effective transfer in low-resource scenarios, and that similar behavior may extend to other low-resource language settings.
title Understanding Cross-Language Transfer Improvements in Low-Resource HTR: The Role of Sequence Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.05900