QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909634760015872 |
|---|---|
| author | Wasfy, Ahmed Nacar, Omer Elkhateb, Abdelakreem Reda, Mahmoud Elshehy, Omar Ammar, Adel Boulila, Wadii |
| author_facet | Wasfy, Ahmed Nacar, Omer Elkhateb, Abdelakreem Reda, Mahmoud Elshehy, Omar Ammar, Adel Boulila, Wadii |
| contents | The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models derived from Qwen2-VL-2B-Instruct, progressively optimized for Arabic through iterative fine-tuning on specialized synthetic datasets. Our leading model, QARI v0.2, establishes a new open-source state-of-the-art with a Word Error Rate (WER) of 0.160, Character Error Rate (CER) of 0.061, and BLEU score of 0.737 on diacritically-rich texts. Qari-OCR demonstrates superior handling of tashkeel, diverse fonts, and document layouts, alongside impressive performance on low-resolution images. Further explorations (QARI v0.3) showcase strong potential for structural document understanding and handwritten text. This work delivers a marked improvement in Arabic OCR accuracy and efficiency, with all models and datasets released to foster further research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_02295 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation Wasfy, Ahmed Nacar, Omer Elkhateb, Abdelakreem Reda, Mahmoud Elshehy, Omar Ammar, Adel Boulila, Wadii Computer Vision and Pattern Recognition Artificial Intelligence The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models derived from Qwen2-VL-2B-Instruct, progressively optimized for Arabic through iterative fine-tuning on specialized synthetic datasets. Our leading model, QARI v0.2, establishes a new open-source state-of-the-art with a Word Error Rate (WER) of 0.160, Character Error Rate (CER) of 0.061, and BLEU score of 0.737 on diacritically-rich texts. Qari-OCR demonstrates superior handling of tashkeel, diverse fonts, and document layouts, alongside impressive performance on low-resolution images. Further explorations (QARI v0.3) showcase strong potential for structural document understanding and handwritten text. This work delivers a marked improvement in Arabic OCR accuracy and efficiency, with all models and datasets released to foster further research. |
| title | QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2506.02295 |