GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hasan, Md. Mahmudul, Choudhury, Ahmed Nesar Tahsin, Hasan, Mahmudul, Khan, Md. Mosaddek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911168878084096
author Hasan, Md. Mahmudul
Choudhury, Ahmed Nesar Tahsin
Hasan, Mahmudul
Khan, Md. Mosaddek
author_facet Hasan, Md. Mahmudul
Choudhury, Ahmed Nesar Tahsin
Hasan, Mahmudul
Khan, Md. Mosaddek
contents Despite Bengali being the sixth most spoken language in the world, handwritten text recognition (HTR) systems for Bengali remain severely underdeveloped. The complexity of Bengali script--featuring conjuncts, diacritics, and highly variable handwriting styles--combined with a scarcity of annotated datasets makes this task particularly challenging. We present GraDeT-HTR, a resource-efficient Bengali handwritten text recognition system based on a Grapheme-aware Decoder-only Transformer architecture. To address the unique challenges of Bengali script, we augment the performance of a decoder-only transformer by integrating a grapheme-based tokenizer and demonstrate that it significantly improves recognition accuracy compared to conventional subword tokenizers. Our model is pretrained on large-scale synthetic data and fine-tuned on real human-annotated samples, achieving state-of-the-art performance on multiple benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer
Hasan, Md. Mahmudul
Choudhury, Ahmed Nesar Tahsin
Hasan, Mahmudul
Khan, Md. Mosaddek
Computer Vision and Pattern Recognition
Despite Bengali being the sixth most spoken language in the world, handwritten text recognition (HTR) systems for Bengali remain severely underdeveloped. The complexity of Bengali script--featuring conjuncts, diacritics, and highly variable handwriting styles--combined with a scarcity of annotated datasets makes this task particularly challenging. We present GraDeT-HTR, a resource-efficient Bengali handwritten text recognition system based on a Grapheme-aware Decoder-only Transformer architecture. To address the unique challenges of Bengali script, we augment the performance of a decoder-only transformer by integrating a grapheme-based tokenizer and demonstrate that it significantly improves recognition accuracy compared to conventional subword tokenizers. Our model is pretrained on large-scale synthetic data and fine-tuned on real human-annotated samples, achieving state-of-the-art performance on multiple benchmark datasets.
title GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.18081