RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915771578318848 |
|---|---|
| author | Köse, Süha Kağan Baytekin, Mehmet Can Aktaş, Burak Görür, Bilge Kaan Munis, Evren Ayberk Yılmaz, Deniz Kartal, Muhammed Yusuf Toraman, Çağrı |
| author_facet | Köse, Süha Kağan Baytekin, Mehmet Can Aktaş, Burak Görür, Bilge Kaan Munis, Evren Ayberk Yılmaz, Deniz Kartal, Muhammed Yusuf Toraman, Çağrı |
| contents | Retrieval-Augmented Generation (RAG) enhances LLM factuality, yet design guidance remains English-centric, limiting insights for morphologically rich languages like Turkish. We address this by constructing a comprehensive Turkish RAG dataset derived from Turkish Wikipedia and CulturaX, comprising question-answer pairs and relevant passage chunks. We benchmark seven stages of the RAG pipeline, from query transformation and reranking to answer refinement, without task-specific fine-tuning. Our results show that complex methods like HyDE maximize accuracy (85%) that is considerably higher than the baseline (78.70%). Also a Pareto-optimal configuration using Cross-encoder Reranking and Context Augmentation achieves comparable performance (84.60%) with much lower cost. We further demonstrate that over-stacking generative modules can degrade performance by distorting morphological cues, whereas simple query clarification with robust reranking offers an effective solution. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03652 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish Köse, Süha Kağan Baytekin, Mehmet Can Aktaş, Burak Görür, Bilge Kaan Munis, Evren Ayberk Yılmaz, Deniz Kartal, Muhammed Yusuf Toraman, Çağrı Computation and Language Artificial Intelligence Information Retrieval Retrieval-Augmented Generation (RAG) enhances LLM factuality, yet design guidance remains English-centric, limiting insights for morphologically rich languages like Turkish. We address this by constructing a comprehensive Turkish RAG dataset derived from Turkish Wikipedia and CulturaX, comprising question-answer pairs and relevant passage chunks. We benchmark seven stages of the RAG pipeline, from query transformation and reranking to answer refinement, without task-specific fine-tuning. Our results show that complex methods like HyDE maximize accuracy (85%) that is considerably higher than the baseline (78.70%). Also a Pareto-optimal configuration using Cross-encoder Reranking and Context Augmentation achieves comparable performance (84.60%) with much lower cost. We further demonstrate that over-stacking generative modules can degrade performance by distorting morphological cues, whereas simple query clarification with robust reranking offers an effective solution. |
| title | RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish |
| topic | Computation and Language Artificial Intelligence Information Retrieval |
| url | https://arxiv.org/abs/2602.03652 |