Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912418101198848 |
|---|---|
| author | Alsubhi, Jumana Alahmadi, Mohammad D. Alhusayni, Ahmed Aldailami, Ibrahim Hamdine, Israa Shabana, Ahmad Iskandar, Yazeed Khayyat, Suhayb |
| author_facet | Alsubhi, Jumana Alahmadi, Mohammad D. Alhusayni, Ahmed Aldailami, Ibrahim Hamdine, Israa Shabana, Ahmad Iskandar, Yazeed Khayyat, Suhayb |
| contents | Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for combining the precision of retrieval systems with the fluency of large language models. While several studies have investigated RAG pipelines for high-resource languages, the optimization of RAG components for Arabic remains underexplored. This study presents a comprehensive empirical evaluation of state-of-the-art RAG components-including chunking strategies, embedding models, rerankers, and language models-across a diverse set of Arabic datasets. Using the RAGAS framework, we systematically compare performance across four core metrics: context precision, context recall, answer faithfulness, and answer relevancy. Our experiments demonstrate that sentence-aware chunking outperforms all other segmentation methods, while BGE-M3 and Multilingual-E5-large emerge as the most effective embedding models. The inclusion of a reranker (bge-reranker-v2-m3) significantly boosts faithfulness in complex datasets, and Aya-8B surpasses StableLM in generation quality. These findings provide critical insights for building high-quality Arabic RAG pipelines and offer practical guidelines for selecting optimal components across different document types. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06339 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components Alsubhi, Jumana Alahmadi, Mohammad D. Alhusayni, Ahmed Aldailami, Ibrahim Hamdine, Israa Shabana, Ahmad Iskandar, Yazeed Khayyat, Suhayb Information Retrieval Artificial Intelligence Computation and Language Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for combining the precision of retrieval systems with the fluency of large language models. While several studies have investigated RAG pipelines for high-resource languages, the optimization of RAG components for Arabic remains underexplored. This study presents a comprehensive empirical evaluation of state-of-the-art RAG components-including chunking strategies, embedding models, rerankers, and language models-across a diverse set of Arabic datasets. Using the RAGAS framework, we systematically compare performance across four core metrics: context precision, context recall, answer faithfulness, and answer relevancy. Our experiments demonstrate that sentence-aware chunking outperforms all other segmentation methods, while BGE-M3 and Multilingual-E5-large emerge as the most effective embedding models. The inclusion of a reranker (bge-reranker-v2-m3) significantly boosts faithfulness in complex datasets, and Aya-8B surpasses StableLM in generation quality. These findings provide critical insights for building high-quality Arabic RAG pipelines and offer practical guidelines for selecting optimal components across different document types. |
| title | Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components |
| topic | Information Retrieval Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2506.06339 |