MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bahaj, Adil, Ghogho, Mounir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912549209899008
author Bahaj, Adil
Ghogho, Mounir
author_facet Bahaj, Adil
Ghogho, Mounir
contents The rapid advancement of large language models (LLMs) has significantly propelled progress in natural language processing (NLP). However, their effectiveness in specialized, low-resource domains-such as Arabic legal contexts-remains limited. This paper introduces MizanQA (pronounced Mizan, meaning "scale" in Arabic, a universal symbol of justice), a benchmark designed to evaluate LLMs on Moroccan legal question answering (QA) tasks, characterised by rich linguistic and legal complexity. The dataset draws on Modern Standard Arabic, Islamic Maliki jurisprudence, Moroccan customary law, and French legal influences. Comprising over 1,700 multiple-choice questions, including multi-answer formats, MizanQA captures the nuances of authentic legal reasoning. Benchmarking experiments with multilingual and Arabic-focused LLMs reveal substantial performance gaps, highlighting the need for tailored evaluation metrics and culturally grounded, domain-specific LLM development.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
Bahaj, Adil
Ghogho, Mounir
Computation and Language
Artificial Intelligence
Information Retrieval
The rapid advancement of large language models (LLMs) has significantly propelled progress in natural language processing (NLP). However, their effectiveness in specialized, low-resource domains-such as Arabic legal contexts-remains limited. This paper introduces MizanQA (pronounced Mizan, meaning "scale" in Arabic, a universal symbol of justice), a benchmark designed to evaluate LLMs on Moroccan legal question answering (QA) tasks, characterised by rich linguistic and legal complexity. The dataset draws on Modern Standard Arabic, Islamic Maliki jurisprudence, Moroccan customary law, and French legal influences. Comprising over 1,700 multiple-choice questions, including multi-answer formats, MizanQA captures the nuances of authentic legal reasoning. Benchmarking experiments with multilingual and Arabic-focused LLMs reveal substantial performance gaps, highlighting the need for tailored evaluation metrics and culturally grounded, domain-specific LLM development.
title MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2508.16357