Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
Fuente:
arXiv
Saved in:
| Main Authors: | AlDahoul, Nouar, Zaki, Yasir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
by: Jan, Essa, et al.
Published: (2024)
by: Jan, Essa, et al.
Published: (2024)
AI-generated faces influence gender stereotypes and racial homogenization
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
by: Aldahoul, Nouar, et al.
Published: (2025)
by: Aldahoul, Nouar, et al.
Published: (2025)
Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
by: Kuo, Chen Wei, et al.
Published: (2025)
by: Kuo, Chen Wei, et al.
Published: (2025)
Inclusive content reduces racial and gender biases, yet non-inclusive content dominates popular culture
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
by: Ibrahim, Hazem, et al.
Published: (2024)
by: Ibrahim, Hazem, et al.
Published: (2024)
Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
by: Hijazi, Faris, et al.
Published: (2024)
by: Hijazi, Faris, et al.
Published: (2024)
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026)
by: Christian, Brian, et al.
Published: (2026)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
by: Ghaboura, Sara, et al.
Published: (2024)
by: Ghaboura, Sara, et al.
Published: (2024)
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
by: Jan, Essa, et al.
Published: (2025)
by: Jan, Essa, et al.
Published: (2025)
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
Fine-tuned Vision Language Model for Localization of Parasitic Eggs in Microscopic Images
by: Sien, Chan Hao, et al.
Published: (2026)
by: Sien, Chan Hao, et al.
Published: (2026)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
by: Marius, Dumitran Adrian, et al.
Published: (2025)
by: Marius, Dumitran Adrian, et al.
Published: (2025)
No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts
by: Fama, Israel, et al.
Published: (2024)
by: Fama, Israel, et al.
Published: (2024)
ALARB: An Arabic Legal Argument Reasoning Benchmark
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
by: Jensen, Benjamin, et al.
Published: (2025)
by: Jensen, Benjamin, et al.
Published: (2025)
Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval
by: Maurya, Anuraj
Published: (2025)
by: Maurya, Anuraj
Published: (2025)
Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models
by: AlDahoul, Nouar, et al.
Published: (2026)
by: AlDahoul, Nouar, et al.
Published: (2026)
LLMs as Writing Assistants: Exploring Perspectives on Sense of Ownership and Reasoning
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
Wikipedia in the Era of LLMs: Evolution and Risks
by: Huang, Siming, et al.
Published: (2025)
by: Huang, Siming, et al.
Published: (2025)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2025)
by: Liu, Haokun, et al.
Published: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
by: Ren, Richard, et al.
Published: (2025)
by: Ren, Richard, et al.
Published: (2025)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
by: Han, Pengrui, et al.
Published: (2025)
by: Han, Pengrui, et al.
Published: (2025)
Similar Items
-
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
by: AlDahoul, Nouar, et al.
Published: (2025) -
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025) -
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024) -
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
by: Jan, Essa, et al.
Published: (2024) -
AI-generated faces influence gender stereotypes and racial homogenization
by: AlDahoul, Nouar, et al.
Published: (2024)