Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hennara, Khalil, Hreden, Muhammad, Hamed, Mohamed Motasim, Bastati, Ahmad, Aldallal, Zeina, Chrouf, Sara, AlModhayan, Safwan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915512731041792
author Hennara, Khalil
Hreden, Muhammad
Hamed, Mohamed Motasim
Bastati, Ahmad
Aldallal, Zeina
Chrouf, Sara
AlModhayan, Safwan
author_facet Hennara, Khalil
Hreden, Muhammad
Hamed, Mohamed Motasim
Bastati, Ahmad
Aldallal, Zeina
Chrouf, Sara
AlModhayan, Safwan
contents Arabic document OCR remains a challenging task due to the language's cursive script, diverse fonts, diacritics, and right-to-left orientation. While modern Multimodal Large Language Models (MLLMs) have advanced document understanding for high-resource languages, their performance on Arabic remains limited. In this work, we introduce Baseer, a vision-language model fine-tuned specifically for Arabic document OCR. Leveraging a large-scale dataset combining synthetic and real-world documents, Baseer is trained using a decoder-only fine-tuning strategy to adapt a pre-trained MLLM while preserving general visual features. We also present Misraj-DocOCR, a high-quality, expert-verified benchmark designed for rigorous evaluation of Arabic OCR systems. Our experiments show that Baseer significantly outperforms existing open-source and commercial solutions, achieving a WER of 0.25 and establishing a new state-of-the-art in the domain of Arabic document OCR. Our results highlight the benefits of domain-specific adaptation of general-purpose MLLMs and establish a strong baseline for high-accuracy OCR on morphologically rich languages like Arabic.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18174
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
Hennara, Khalil
Hreden, Muhammad
Hamed, Mohamed Motasim
Bastati, Ahmad
Aldallal, Zeina
Chrouf, Sara
AlModhayan, Safwan
Computer Vision and Pattern Recognition
Computation and Language
Arabic document OCR remains a challenging task due to the language's cursive script, diverse fonts, diacritics, and right-to-left orientation. While modern Multimodal Large Language Models (MLLMs) have advanced document understanding for high-resource languages, their performance on Arabic remains limited. In this work, we introduce Baseer, a vision-language model fine-tuned specifically for Arabic document OCR. Leveraging a large-scale dataset combining synthetic and real-world documents, Baseer is trained using a decoder-only fine-tuning strategy to adapt a pre-trained MLLM while preserving general visual features. We also present Misraj-DocOCR, a high-quality, expert-verified benchmark designed for rigorous evaluation of Arabic OCR systems. Our experiments show that Baseer significantly outperforms existing open-source and commercial solutions, achieving a WER of 0.25 and establishing a new state-of-the-art in the domain of Arabic document OCR. Our results highlight the benefits of domain-specific adaptation of general-purpose MLLMs and establish a strong baseline for high-accuracy OCR on morphologically rich languages like Arabic.
title Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.18174