QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wasfy, Ahmed, Nacar, Omer, Elkhateb, Abdelakreem, Reda, Mahmoud, Elshehy, Omar, Ammar, Adel, Boulila, Wadii
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909634760015872
author Wasfy, Ahmed
Nacar, Omer
Elkhateb, Abdelakreem
Reda, Mahmoud
Elshehy, Omar
Ammar, Adel
Boulila, Wadii
author_facet Wasfy, Ahmed
Nacar, Omer
Elkhateb, Abdelakreem
Reda, Mahmoud
Elshehy, Omar
Ammar, Adel
Boulila, Wadii
contents The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models derived from Qwen2-VL-2B-Instruct, progressively optimized for Arabic through iterative fine-tuning on specialized synthetic datasets. Our leading model, QARI v0.2, establishes a new open-source state-of-the-art with a Word Error Rate (WER) of 0.160, Character Error Rate (CER) of 0.061, and BLEU score of 0.737 on diacritically-rich texts. Qari-OCR demonstrates superior handling of tashkeel, diverse fonts, and document layouts, alongside impressive performance on low-resolution images. Further explorations (QARI v0.3) showcase strong potential for structural document understanding and handwritten text. This work delivers a marked improvement in Arabic OCR accuracy and efficiency, with all models and datasets released to foster further research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02295
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
Wasfy, Ahmed
Nacar, Omer
Elkhateb, Abdelakreem
Reda, Mahmoud
Elshehy, Omar
Ammar, Adel
Boulila, Wadii
Computer Vision and Pattern Recognition
Artificial Intelligence
The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models derived from Qwen2-VL-2B-Instruct, progressively optimized for Arabic through iterative fine-tuning on specialized synthetic datasets. Our leading model, QARI v0.2, establishes a new open-source state-of-the-art with a Word Error Rate (WER) of 0.160, Character Error Rate (CER) of 0.061, and BLEU score of 0.737 on diacritically-rich texts. Qari-OCR demonstrates superior handling of tashkeel, diverse fonts, and document layouts, alongside impressive performance on low-resolution images. Further explorations (QARI v0.3) showcase strong potential for structural document understanding and handwritten text. This work delivers a marked improvement in Arabic OCR accuracy and efficiency, with all models and datasets released to foster further research.
title QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.02295