RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918394400342016 |
|---|---|
| author | Cohen, Dvir Houri, Tamir Burg, Lin Barkan, Gilad |
| author_facet | Cohen, Dvir Houri, Tamir Burg, Lin Barkan, Gilad |
| contents | Retrieval-Augmented Generation (RAG) systems couple large language models with external knowledge, yet most evaluation methods report aggregate scores that reveal whether a pipeline underperforms but not where or why. We introduce RAGXplain, an evaluation framework that translates performance metrics into actionable guidance. RAGXplain structures evaluation around a 'Metric Diamond' connecting user input, retrieved context, generated answer, and (when available) ground truth via six diagnostic dimensions. It uses LLM reasoning to produce natural-language failure-mode explanations and prioritized interventions. Across five QA benchmarks, applying RAGXplain's recommendations in a single human-guided pass consistently improves RAG pipeline performance across multiple metrics. We release RAGXplain as open source to support reproducibility and community adoption. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13538 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines Cohen, Dvir Houri, Tamir Burg, Lin Barkan, Gilad Information Retrieval Artificial Intelligence Retrieval-Augmented Generation (RAG) systems couple large language models with external knowledge, yet most evaluation methods report aggregate scores that reveal whether a pipeline underperforms but not where or why. We introduce RAGXplain, an evaluation framework that translates performance metrics into actionable guidance. RAGXplain structures evaluation around a 'Metric Diamond' connecting user input, retrieved context, generated answer, and (when available) ground truth via six diagnostic dimensions. It uses LLM reasoning to produce natural-language failure-mode explanations and prioritized interventions. Across five QA benchmarks, applying RAGXplain's recommendations in a single human-guided pass consistently improves RAG pipeline performance across multiple metrics. We release RAGXplain as open source to support reproducibility and community adoption. |
| title | RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines |
| topic | Information Retrieval Artificial Intelligence |
| url | https://arxiv.org/abs/2505.13538 |