Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bouchekif, Abdessalam, Rashwani, Samer, Sbahi, Heba, Gaben, Shahd, Al-Khatib, Mutaz, Ghaly, Mohammed
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916954232586240
author Bouchekif, Abdessalam
Rashwani, Samer
Sbahi, Heba
Gaben, Shahd
Al-Khatib, Mutaz
Ghaly, Mohammed
author_facet Bouchekif, Abdessalam
Rashwani, Samer
Sbahi, Heba
Gaben, Shahd
Al-Khatib, Mutaz
Ghaly, Mohammed
contents This paper evaluates the knowledge and reasoning capabilities of Large Language Models in Islamic inheritance law, known as 'ilm al-mawarith. We assess the performance of seven LLMs using a benchmark of 1,000 multiple-choice questions covering diverse inheritance scenarios, designed to test models' ability to understand the inheritance context and compute the distribution of shares prescribed by Islamic jurisprudence. The results reveal a significant performance gap: o3 and Gemini 2.5 achieved accuracies above 90%, whereas ALLaM, Fanar, LLaMA, and Mistral scored below 50%. These disparities reflect important differences in reasoning ability and domain adaptation. We conduct a detailed error analysis to identify recurring failure patterns across models, including misunderstandings of inheritance scenarios, incorrect application of legal rules, and insufficient domain knowledge. Our findings highlight limitations in handling structured legal reasoning and suggest directions for improving performance in Islamic legal reasoning. Code: https://github.com/bouchekif/inheritance_evaluation
format Preprint
id arxiv_https___arxiv_org_abs_2509_01081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
Bouchekif, Abdessalam
Rashwani, Samer
Sbahi, Heba
Gaben, Shahd
Al-Khatib, Mutaz
Ghaly, Mohammed
Computation and Language
Artificial Intelligence
I.2.6; I.2.7
This paper evaluates the knowledge and reasoning capabilities of Large Language Models in Islamic inheritance law, known as 'ilm al-mawarith. We assess the performance of seven LLMs using a benchmark of 1,000 multiple-choice questions covering diverse inheritance scenarios, designed to test models' ability to understand the inheritance context and compute the distribution of shares prescribed by Islamic jurisprudence. The results reveal a significant performance gap: o3 and Gemini 2.5 achieved accuracies above 90%, whereas ALLaM, Fanar, LLaMA, and Mistral scored below 50%. These disparities reflect important differences in reasoning ability and domain adaptation. We conduct a detailed error analysis to identify recurring failure patterns across models, including misunderstandings of inheritance scenarios, incorrect application of legal rules, and insufficient domain knowledge. Our findings highlight limitations in handling structured legal reasoning and suggest directions for improving performance in Islamic legal reasoning. Code: https://github.com/bouchekif/inheritance_evaluation
title Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
topic Computation and Language
Artificial Intelligence
I.2.6; I.2.7
url https://arxiv.org/abs/2509.01081