Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bhatti, Hunzalah Hassan, Alam, Firoj
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911599714893824
author Bhatti, Hunzalah Hassan
Alam, Firoj
author_facet Bhatti, Hunzalah Hassan
Alam, Firoj
contents Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages. We propose a comprehensive method that (i) translates Modern Standard Arabic (MSA) multiple-choice questions (MCQs) into English and several Arabic dialects, (ii) converts them into open-ended questions (OEQs), (iii) benchmarks a range of zero-shot and fine-tuned LLMs under both MCQ and OEQ settings, and (iv) generates chain-of-thought (CoT) rationales to fine-tune models for step-by-step reasoning. Using this method, we extend an existing dataset in which QAs are parallelly aligned across multiple language varieties, making it, to our knowledge, the first of its kind. We conduct extensive experiments with both open and closed models. Our findings show that (i) models underperform on Arabic dialects, revealing persistent gaps in culturally grounded and dialect-specific knowledge; (ii) Arabic-centric models perform well on MCQs but struggle with OEQs; and (iii) CoT improves judged correctness while yielding mixed n-gram-based metrics. The developed dataset will be publicly released to support further research on culturally and linguistically inclusive evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
Bhatti, Hunzalah Hassan
Alam, Firoj
Computation and Language
Artificial Intelligence
68T50
F.2.2; I.2.7
Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages. We propose a comprehensive method that (i) translates Modern Standard Arabic (MSA) multiple-choice questions (MCQs) into English and several Arabic dialects, (ii) converts them into open-ended questions (OEQs), (iii) benchmarks a range of zero-shot and fine-tuned LLMs under both MCQ and OEQ settings, and (iv) generates chain-of-thought (CoT) rationales to fine-tune models for step-by-step reasoning. Using this method, we extend an existing dataset in which QAs are parallelly aligned across multiple language varieties, making it, to our knowledge, the first of its kind. We conduct extensive experiments with both open and closed models. Our findings show that (i) models underperform on Arabic dialects, revealing persistent gaps in culturally grounded and dialect-specific knowledge; (ii) Arabic-centric models perform well on MCQs but struggle with OEQs; and (iii) CoT improves judged correctness while yielding mixed n-gram-based metrics. The developed dataset will be publicly released to support further research on culturally and linguistically inclusive evaluation.
title Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
topic Computation and Language
Artificial Intelligence
68T50
F.2.2; I.2.7
url https://arxiv.org/abs/2510.24328