Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Basem, Mohamed, Oshallah, Islam, Hamdi, Ali, Shaban, Khaled, Kassab, Hozaifa
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918135187111936
author Basem, Mohamed
Oshallah, Islam
Hamdi, Ali
Shaban, Khaled
Kassab, Hozaifa
author_facet Basem, Mohamed
Oshallah, Islam
Hamdi, Ali
Shaban, Khaled
Kassab, Hozaifa
contents Quranic Question Answering presents unique challenges due to the linguistic complexity of Classical Arabic and the semantic richness of religious texts. In this paper, we propose a novel two-stage framework that addresses both passage retrieval and answer extraction. For passage retrieval, we ensemble fine-tuned Arabic language models to achieve superior ranking performance. For answer extraction, we employ instruction-tuned large language models with few-shot prompting to overcome the limitations of fine-tuning on small datasets. Our approach achieves state-of-the-art results on the Quran QA 2023 Shared Task, with a MAP@10 of 0.3128 and MRR@10 of 0.5763 for retrieval, and a pAP@10 of 0.669 for extraction, substantially outperforming previous methods. These results demonstrate that combining model ensembling and instruction-tuned language models effectively addresses the challenges of low-resource question answering in specialized domains.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06971
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
Basem, Mohamed
Oshallah, Islam
Hamdi, Ali
Shaban, Khaled
Kassab, Hozaifa
Computation and Language
Information Retrieval
Quranic Question Answering presents unique challenges due to the linguistic complexity of Classical Arabic and the semantic richness of religious texts. In this paper, we propose a novel two-stage framework that addresses both passage retrieval and answer extraction. For passage retrieval, we ensemble fine-tuned Arabic language models to achieve superior ranking performance. For answer extraction, we employ instruction-tuned large language models with few-shot prompting to overcome the limitations of fine-tuning on small datasets. Our approach achieves state-of-the-art results on the Quran QA 2023 Shared Task, with a MAP@10 of 0.3128 and MRR@10 of 0.5763 for retrieval, and a pAP@10 of 0.669 for extraction, substantially outperforming previous methods. These results demonstrate that combining model ensembling and instruction-tuned language models effectively addresses the challenges of low-resource question answering in specialized domains.
title Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2508.06971