IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Abdelaal, Ali, Haffar, Mohammed Nader Al, Fawzi, Mahmoud, Magdy, Walid |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
by: Elmahjub, Ezieddin, et al.
Published: (2026)
by: Elmahjub, Ezieddin, et al.
Published: (2026)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
From RAG to Agentic RAG for Faithful Islamic Question Answering
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
by: Pramodya, Ashmari, et al.
Published: (2025)
by: Pramodya, Ashmari, et al.
Published: (2025)
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents
by: Sayeed, Mohammad Amaan, et al.
Published: (2025)
by: Sayeed, Mohammad Amaan, et al.
Published: (2025)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
by: Togmanov, Mukhammed, et al.
Published: (2025)
by: Togmanov, Mukhammed, et al.
Published: (2025)
Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA
by: Abbas, Ummar, et al.
Published: (2026)
by: Abbas, Ummar, et al.
Published: (2026)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
by: Bouchekif, Abdessalam, et al.
Published: (2025)
by: Bouchekif, Abdessalam, et al.
Published: (2025)
Building an Efficient Multilingual Non-Profit IR System for the Islamic Domain Leveraging Multiprocessing Design in Rust
by: Pavlova, Vera, et al.
Published: (2024)
by: Pavlova, Vera, et al.
Published: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
by: Plaza, Irene, et al.
Published: (2024)
by: Plaza, Irene, et al.
Published: (2024)
LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama
by: Etori, Naome A., et al.
Published: (2025)
by: Etori, Naome A., et al.
Published: (2025)
Fabricating Holiness: Characterizing Religious Misinformation Circulators on Arabic Social Media
by: Fawzi, Mahmoud, et al.
Published: (2025)
by: Fawzi, Mahmoud, et al.
Published: (2025)
"The Prophet said so!": On Exploring Hadith Presence on Arabic Social Media
by: Fawzi, Mahmoud, et al.
Published: (2024)
by: Fawzi, Mahmoud, et al.
Published: (2024)
Contemporary Bioethics Islamic Perspective
by: Mohammed Ali Al-Bar
by: Mohammed Ali Al-Bar
A Benchmark Dataset with Larger Context for Non-Factoid Question Answering over Islamic Text
by: Qamar, Faiza, et al.
Published: (2024)
by: Qamar, Faiza, et al.
Published: (2024)
Revisiting Common Assumptions about Arabic Dialects in NLP
by: Keleg, Amr, et al.
Published: (2025)
by: Keleg, Amr, et al.
Published: (2025)
Culture Matters in Toxic Language Detection in Persian
by: Bokaei, Zahra, et al.
Published: (2025)
by: Bokaei, Zahra, et al.
Published: (2025)
Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
by: Pavlova, Vera, et al.
Published: (2025)
by: Pavlova, Vera, et al.
Published: (2025)
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
Are We Done with MMLU?
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Transformer Tafsir at QIAS 2025 Shared Task: Hybrid Retrieval-Augmented Generation for Islamic Knowledge Question Answering
by: Ahmad, Muhammad Abu, et al.
Published: (2025)
by: Ahmad, Muhammad Abu, et al.
Published: (2025)
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
by: Singh, Shivalika, et al.
Published: (2024)
by: Singh, Shivalika, et al.
Published: (2024)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
by: Bayram, M. Ali, et al.
Published: (2024)
by: Bayram, M. Ali, et al.
Published: (2024)
Multi-stage Training of Bilingual Islamic LLM for Neural Passage Retrieval
by: Pavlova, Vera
Published: (2025)
by: Pavlova, Vera
Published: (2025)
Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets
by: Keleg, Amr, et al.
Published: (2024)
by: Keleg, Amr, et al.
Published: (2024)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
by: Mustapha, Ahmad, et al.
Published: (2024)
by: Mustapha, Ahmad, et al.
Published: (2024)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
by: Wang, Wentian, et al.
Published: (2024)
by: Wang, Wentian, et al.
Published: (2024)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
by: KJ, Sankalp, et al.
Published: (2025)
by: KJ, Sankalp, et al.
Published: (2025)
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
by: Yüksel, Arda, et al.
Published: (2024)
by: Yüksel, Arda, et al.
Published: (2024)
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
by: Joy, Saman Sarker, et al.
Published: (2025)
by: Joy, Saman Sarker, et al.
Published: (2025)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Similar Items
-
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
by: Elmahjub, Ezieddin, et al.
Published: (2026) -
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025) -
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
by: Mushtaq, Abdullah, et al.
Published: (2025) -
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025) -
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)