ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghaboura, Sara, More, Ketan, Alghallabi, Wafa, Thawakar, Omkar, Laaksonen, Jorma, Cholakkal, Hisham, Khan, Salman, Anwer, Rao Muhammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918031007940608
author Ghaboura, Sara
More, Ketan
Alghallabi, Wafa
Thawakar, Omkar
Laaksonen, Jorma
Cholakkal, Hisham
Khan, Salman
Anwer, Rao Muhammad
author_facet Ghaboura, Sara
More, Ketan
Alghallabi, Wafa
Thawakar, Omkar
Laaksonen, Jorma
Cholakkal, Hisham
Khan, Salman
Anwer, Rao Muhammad
contents As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks remain focused on English, overlooking languages with rich linguistic and cultural contexts, such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reasoning Benchmark (ARB), the first benchmark designed to evaluate step-by-step reasoning in Arabic across both textual and visual modalities. ARB spans 11 diverse domains, including visual reasoning, document understanding, OCR, scientific analysis, and cultural interpretation. It comprises 1,356 multimodal samples paired with 5,119 human-curated reasoning steps and corresponding actions. We evaluated 12 state-of-the-art open- and closed-source LMMs and found persistent challenges in coherence, faithfulness, and cultural grounding. ARB offers a structured framework for diagnosing multimodal reasoning in underrepresented languages and marks a critical step toward inclusive, transparent, and culturally aware AI systems. We release the benchmark, rubric, and evaluation suit to support future research and reproducibility. Code available at: https://github.com/mbzuai-oryx/ARB
format Preprint
id arxiv_https___arxiv_org_abs_2505_17021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
Ghaboura, Sara
More, Ketan
Alghallabi, Wafa
Thawakar, Omkar
Laaksonen, Jorma
Cholakkal, Hisham
Khan, Salman
Anwer, Rao Muhammad
Computer Vision and Pattern Recognition
As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks remain focused on English, overlooking languages with rich linguistic and cultural contexts, such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reasoning Benchmark (ARB), the first benchmark designed to evaluate step-by-step reasoning in Arabic across both textual and visual modalities. ARB spans 11 diverse domains, including visual reasoning, document understanding, OCR, scientific analysis, and cultural interpretation. It comprises 1,356 multimodal samples paired with 5,119 human-curated reasoning steps and corresponding actions. We evaluated 12 state-of-the-art open- and closed-source LMMs and found persistent challenges in coherence, faithfulness, and cultural grounding. ARB offers a structured framework for diagnosing multimodal reasoning in underrepresented languages and marks a critical step toward inclusive, transparent, and culturally aware AI systems. We release the benchmark, rubric, and evaluation suit to support future research and reproducibility. Code available at: https://github.com/mbzuai-oryx/ARB
title ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17021