Muharaf: Manuscripts of Handwritten Arabic Dataset for Cursive Text Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saeed, Mehreen, Chan, Adrian, Mijar, Anupam, Moukarzel, Joseph, Habchi, Georges, Younes, Carlos, Elias, Amin, Wong, Chau-Wai, Khater, Akram
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912220148924416
author Saeed, Mehreen
Chan, Adrian
Mijar, Anupam
Moukarzel, Joseph
Habchi, Georges
Younes, Carlos
Elias, Amin
Wong, Chau-Wai
Khater, Akram
author_facet Saeed, Mehreen
Chan, Adrian
Mijar, Anupam
Moukarzel, Joseph
Habchi, Georges
Younes, Carlos
Elias, Amin
Wong, Chau-Wai
Khater, Akram
contents We present the Manuscripts of Handwritten Arabic~(Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is accompanied by spatial polygonal coordinates of its text lines as well as basic page elements. This dataset was compiled to advance the state of the art in handwritten text recognition (HTR), not only for Arabic manuscripts but also for cursive text in general. The Muharaf dataset includes diverse handwriting styles and a wide range of document types, including personal letters, diaries, notes, poems, church records, and legal correspondences. In this paper, we describe the data acquisition pipeline, notable dataset features, and statistics. We also provide a preliminary baseline result achieved by training convolutional neural networks using this data.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09630
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Muharaf: Manuscripts of Handwritten Arabic Dataset for Cursive Text Recognition
Saeed, Mehreen
Chan, Adrian
Mijar, Anupam
Moukarzel, Joseph
Habchi, Georges
Younes, Carlos
Elias, Amin
Wong, Chau-Wai
Khater, Akram
Computer Vision and Pattern Recognition
Machine Learning
We present the Manuscripts of Handwritten Arabic~(Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is accompanied by spatial polygonal coordinates of its text lines as well as basic page elements. This dataset was compiled to advance the state of the art in handwritten text recognition (HTR), not only for Arabic manuscripts but also for cursive text in general. The Muharaf dataset includes diverse handwriting styles and a wide range of document types, including personal letters, diaries, notes, poems, church records, and legal correspondences. In this paper, we describe the data acquisition pipeline, notable dataset features, and statistics. We also provide a preliminary baseline result achieved by training convolutional neural networks using this data.
title Muharaf: Manuscripts of Handwritten Arabic Dataset for Cursive Text Recognition
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.09630