DRAGOn: Designing RAG On Periodically Updated Corpus

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chernogorskii, Fedor, Averkiev, Sergei, Kudraleeva, Liliya, Martirosian, Zaven, Tikhonova, Maria, Malykh, Valentin, Fenogenova, Alena
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910015922634752
author Chernogorskii, Fedor
Averkiev, Sergei
Kudraleeva, Liliya
Martirosian, Zaven
Tikhonova, Maria
Malykh, Valentin
Fenogenova, Alena
author_facet Chernogorskii, Fedor
Averkiev, Sergei
Kudraleeva, Liliya
Martirosian, Zaven
Tikhonova, Maria
Malykh, Valentin
Fenogenova, Alena
contents This paper introduces DRAGOn, method to design a RAG benchmark on a regularly updated corpus. It features recent reference datasets, a question generation framework, an automatic evaluation pipeline, and a public leaderboard. Specified reference datasets allow for uniform comparison of RAG systems, while newly generated dataset versions mitigate data leakage and ensure that all models are evaluated on unseen, comparable data. The pipeline for automatic question generation extracts the Knowledge Graph from the text corpus and produces multiple question-answer pairs utilizing modern LLM capabilities. A set of diverse LLM-as-Judge metrics is provided for a comprehensive model evaluation. We used Russian news outlets to form the datasets and demonstrate our methodology. We launch a public leaderboard to track the development of RAG systems and encourage community participation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DRAGOn: Designing RAG On Periodically Updated Corpus
Chernogorskii, Fedor
Averkiev, Sergei
Kudraleeva, Liliya
Martirosian, Zaven
Tikhonova, Maria
Malykh, Valentin
Fenogenova, Alena
Computation and Language
Artificial Intelligence
This paper introduces DRAGOn, method to design a RAG benchmark on a regularly updated corpus. It features recent reference datasets, a question generation framework, an automatic evaluation pipeline, and a public leaderboard. Specified reference datasets allow for uniform comparison of RAG systems, while newly generated dataset versions mitigate data leakage and ensure that all models are evaluated on unseen, comparable data. The pipeline for automatic question generation extracts the Knowledge Graph from the text corpus and produces multiple question-answer pairs utilizing modern LLM capabilities. A set of diverse LLM-as-Judge metrics is provided for a comprehensive model evaluation. We used Russian news outlets to form the datasets and demonstrate our methodology. We launch a public leaderboard to track the development of RAG systems and encourage community participation.
title DRAGOn: Designing RAG On Periodically Updated Corpus
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.05713