RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Dvir, Houri, Tamir, Burg, Lin, Barkan, Gilad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918394400342016
author Cohen, Dvir
Houri, Tamir
Burg, Lin
Barkan, Gilad
author_facet Cohen, Dvir
Houri, Tamir
Burg, Lin
Barkan, Gilad
contents Retrieval-Augmented Generation (RAG) systems couple large language models with external knowledge, yet most evaluation methods report aggregate scores that reveal whether a pipeline underperforms but not where or why. We introduce RAGXplain, an evaluation framework that translates performance metrics into actionable guidance. RAGXplain structures evaluation around a 'Metric Diamond' connecting user input, retrieved context, generated answer, and (when available) ground truth via six diagnostic dimensions. It uses LLM reasoning to produce natural-language failure-mode explanations and prioritized interventions. Across five QA benchmarks, applying RAGXplain's recommendations in a single human-guided pass consistently improves RAG pipeline performance across multiple metrics. We release RAGXplain as open source to support reproducibility and community adoption.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines
Cohen, Dvir
Houri, Tamir
Burg, Lin
Barkan, Gilad
Information Retrieval
Artificial Intelligence
Retrieval-Augmented Generation (RAG) systems couple large language models with external knowledge, yet most evaluation methods report aggregate scores that reveal whether a pipeline underperforms but not where or why. We introduce RAGXplain, an evaluation framework that translates performance metrics into actionable guidance. RAGXplain structures evaluation around a 'Metric Diamond' connecting user input, retrieved context, generated answer, and (when available) ground truth via six diagnostic dimensions. It uses LLM reasoning to produce natural-language failure-mode explanations and prioritized interventions. Across five QA benchmarks, applying RAGXplain's recommendations in a single human-guided pass consistently improves RAG pipeline performance across multiple metrics. We release RAGXplain as open source to support reproducibility and community adoption.
title RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2505.13538