Can we Evaluate RAGs with Synthetic Data?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: van Elburg, Jonas, van der Putten, Peter, Marx, Maarten
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909859897671680
author van Elburg, Jonas
van der Putten, Peter
Marx, Maarten
author_facet van Elburg, Jonas
van der Putten, Peter
Marx, Maarten
contents We investigate whether synthetic question-answer (QA) data generated by large language models (LLMs) can serve as an effective proxy for human-labeled benchmarks when the latter is unavailable. We assess the reliability of synthetic benchmarks across two experiments: one varying retriever parameters while keeping the generator fixed, and another varying the generator with fixed retriever parameters. Across four datasets, of which two open-domain and two proprietary, we find that synthetic benchmarks reliably rank the RAGs varying in terms of retriever configuration, aligning well with human-labeled benchmark baselines. However, they do not consistently produce reliable RAG rankings when comparing generator architectures. The breakdown possibly arises from a combination of task mismatch between the synthetic and human benchmarks, and stylistic bias favoring certain generators.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can we Evaluate RAGs with Synthetic Data?
van Elburg, Jonas
van der Putten, Peter
Marx, Maarten
Computation and Language
Artificial Intelligence
We investigate whether synthetic question-answer (QA) data generated by large language models (LLMs) can serve as an effective proxy for human-labeled benchmarks when the latter is unavailable. We assess the reliability of synthetic benchmarks across two experiments: one varying retriever parameters while keeping the generator fixed, and another varying the generator with fixed retriever parameters. Across four datasets, of which two open-domain and two proprietary, we find that synthetic benchmarks reliably rank the RAGs varying in terms of retriever configuration, aligning well with human-labeled benchmark baselines. However, they do not consistently produce reliable RAG rankings when comparing generator architectures. The breakdown possibly arises from a combination of task mismatch between the synthetic and human benchmarks, and stylistic bias favoring certain generators.
title Can we Evaluate RAGs with Synthetic Data?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.11758