RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Köse, Süha Kağan, Baytekin, Mehmet Can, Aktaş, Burak, Görür, Bilge Kaan, Munis, Evren Ayberk, Yılmaz, Deniz, Kartal, Muhammed Yusuf, Toraman, Çağrı
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915771578318848
author Köse, Süha Kağan
Baytekin, Mehmet Can
Aktaş, Burak
Görür, Bilge Kaan
Munis, Evren Ayberk
Yılmaz, Deniz
Kartal, Muhammed Yusuf
Toraman, Çağrı
author_facet Köse, Süha Kağan
Baytekin, Mehmet Can
Aktaş, Burak
Görür, Bilge Kaan
Munis, Evren Ayberk
Yılmaz, Deniz
Kartal, Muhammed Yusuf
Toraman, Çağrı
contents Retrieval-Augmented Generation (RAG) enhances LLM factuality, yet design guidance remains English-centric, limiting insights for morphologically rich languages like Turkish. We address this by constructing a comprehensive Turkish RAG dataset derived from Turkish Wikipedia and CulturaX, comprising question-answer pairs and relevant passage chunks. We benchmark seven stages of the RAG pipeline, from query transformation and reranking to answer refinement, without task-specific fine-tuning. Our results show that complex methods like HyDE maximize accuracy (85%) that is considerably higher than the baseline (78.70%). Also a Pareto-optimal configuration using Cross-encoder Reranking and Context Augmentation achieves comparable performance (84.60%) with much lower cost. We further demonstrate that over-stacking generative modules can degrade performance by distorting morphological cues, whereas simple query clarification with robust reranking offers an effective solution.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03652
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
Köse, Süha Kağan
Baytekin, Mehmet Can
Aktaş, Burak
Görür, Bilge Kaan
Munis, Evren Ayberk
Yılmaz, Deniz
Kartal, Muhammed Yusuf
Toraman, Çağrı
Computation and Language
Artificial Intelligence
Information Retrieval
Retrieval-Augmented Generation (RAG) enhances LLM factuality, yet design guidance remains English-centric, limiting insights for morphologically rich languages like Turkish. We address this by constructing a comprehensive Turkish RAG dataset derived from Turkish Wikipedia and CulturaX, comprising question-answer pairs and relevant passage chunks. We benchmark seven stages of the RAG pipeline, from query transformation and reranking to answer refinement, without task-specific fine-tuning. Our results show that complex methods like HyDE maximize accuracy (85%) that is considerably higher than the baseline (78.70%). Also a Pareto-optimal configuration using Cross-encoder Reranking and Context Augmentation achieves comparable performance (84.60%) with much lower cost. We further demonstrate that over-stacking generative modules can degrade performance by distorting morphological cues, whereas simple query clarification with robust reranking offers an effective solution.
title RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2602.03652