Towards a rigorous evaluation of RAG systems: the challenge of due diligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Martinon, Grégoire, de Brionne, Alexandra Lorenzo, Bohard, Jérôme, Lojou, Antoine, Hervault, Damien, Brunel, Nicolas J-B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916868778885120
author Martinon, Grégoire
de Brionne, Alexandra Lorenzo
Bohard, Jérôme
Lojou, Antoine
Hervault, Damien
Brunel, Nicolas J-B.
author_facet Martinon, Grégoire
de Brionne, Alexandra Lorenzo
Bohard, Jérôme
Lojou, Antoine
Hervault, Damien
Brunel, Nicolas J-B.
contents The rise of generative AI, has driven significant advancements in high-risk sectors like healthcare and finance. The Retrieval-Augmented Generation (RAG) architecture, combining language models (LLMs) with search engines, is particularly notable for its ability to generate responses from document corpora. Despite its potential, the reliability of RAG systems in critical contexts remains a concern, with issues such as hallucinations persisting. This study evaluates a RAG system used in due diligence for an investment fund. We propose a robust evaluation protocol combining human annotations and LLM-Judge annotations to identify system failures, like hallucinations, off-topic, failed citations, and abstentions. Inspired by the Prediction Powered Inference (PPI) method, we achieve precise performance measurements with statistical guarantees. We provide a comprehensive dataset for further analysis. Our contributions aim to enhance the reliability and scalability of RAG systems evaluation protocols in industrial applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards a rigorous evaluation of RAG systems: the challenge of due diligence
Martinon, Grégoire
de Brionne, Alexandra Lorenzo
Bohard, Jérôme
Lojou, Antoine
Hervault, Damien
Brunel, Nicolas J-B.
Artificial Intelligence
Applications
The rise of generative AI, has driven significant advancements in high-risk sectors like healthcare and finance. The Retrieval-Augmented Generation (RAG) architecture, combining language models (LLMs) with search engines, is particularly notable for its ability to generate responses from document corpora. Despite its potential, the reliability of RAG systems in critical contexts remains a concern, with issues such as hallucinations persisting. This study evaluates a RAG system used in due diligence for an investment fund. We propose a robust evaluation protocol combining human annotations and LLM-Judge annotations to identify system failures, like hallucinations, off-topic, failed citations, and abstentions. Inspired by the Prediction Powered Inference (PPI) method, we achieve precise performance measurements with statistical guarantees. We provide a comprehensive dataset for further analysis. Our contributions aim to enhance the reliability and scalability of RAG systems evaluation protocols in industrial applications.
title Towards a rigorous evaluation of RAG systems: the challenge of due diligence
topic Artificial Intelligence
Applications
url https://arxiv.org/abs/2507.21753