Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918135187111936 |
|---|---|
| author | Basem, Mohamed Oshallah, Islam Hamdi, Ali Shaban, Khaled Kassab, Hozaifa |
| author_facet | Basem, Mohamed Oshallah, Islam Hamdi, Ali Shaban, Khaled Kassab, Hozaifa |
| contents | Quranic Question Answering presents unique challenges due to the linguistic complexity of Classical Arabic and the semantic richness of religious texts. In this paper, we propose a novel two-stage framework that addresses both passage retrieval and answer extraction. For passage retrieval, we ensemble fine-tuned Arabic language models to achieve superior ranking performance. For answer extraction, we employ instruction-tuned large language models with few-shot prompting to overcome the limitations of fine-tuning on small datasets. Our approach achieves state-of-the-art results on the Quran QA 2023 Shared Task, with a MAP@10 of 0.3128 and MRR@10 of 0.5763 for retrieval, and a pAP@10 of 0.669 for extraction, substantially outperforming previous methods. These results demonstrate that combining model ensembling and instruction-tuned language models effectively addresses the challenges of low-resource question answering in specialized domains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_06971 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction Basem, Mohamed Oshallah, Islam Hamdi, Ali Shaban, Khaled Kassab, Hozaifa Computation and Language Information Retrieval Quranic Question Answering presents unique challenges due to the linguistic complexity of Classical Arabic and the semantic richness of religious texts. In this paper, we propose a novel two-stage framework that addresses both passage retrieval and answer extraction. For passage retrieval, we ensemble fine-tuned Arabic language models to achieve superior ranking performance. For answer extraction, we employ instruction-tuned large language models with few-shot prompting to overcome the limitations of fine-tuning on small datasets. Our approach achieves state-of-the-art results on the Quran QA 2023 Shared Task, with a MAP@10 of 0.3128 and MRR@10 of 0.5763 for retrieval, and a pAP@10 of 0.669 for extraction, substantially outperforming previous methods. These results demonstrate that combining model ensembling and instruction-tuned language models effectively addresses the challenges of low-resource question answering in specialized domains. |
| title | Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction |
| topic | Computation and Language Information Retrieval |
| url | https://arxiv.org/abs/2508.06971 |