Source Attribution in Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nematov, Ikhtiyor, Kalai, Tarik, Kuzmenko, Elizaveta, Fugagnoli, Gabriele, Sacharidis, Dimitris, Hose, Katja, Sagi, Tomer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913929577365504
author Nematov, Ikhtiyor
Kalai, Tarik
Kuzmenko, Elizaveta
Fugagnoli, Gabriele
Sacharidis, Dimitris
Hose, Katja
Sagi, Tomer
author_facet Nematov, Ikhtiyor
Kalai, Tarik
Kuzmenko, Elizaveta
Fugagnoli, Gabriele
Sacharidis, Dimitris
Hose, Katja
Sagi, Tomer
contents While attribution methods, such as Shapley values, are widely used to explain the importance of features or training data in traditional machine learning, their application to Large Language Models (LLMs), particularly within Retrieval-Augmented Generation (RAG) systems, is nascent and challenging. The primary obstacle is the substantial computational cost, where each utility function evaluation involves an expensive LLM call, resulting in direct monetary and time expenses. This paper investigates the feasibility and effectiveness of adapting Shapley-based attribution to identify influential retrieved documents in RAG. We compare Shapley with more computationally tractable approximations and some existing attribution methods for LLM. Our work aims to: (1) systematically apply established attribution principles to the RAG document-level setting; (2) quantify how well SHAP approximations can mirror exact attributions while minimizing costly LLM interactions; and (3) evaluate their practical explainability in identifying critical documents, especially under complex inter-document relationships such as redundancy, complementarity, and synergy. This study seeks to bridge the gap between powerful attribution techniques and the practical constraints of LLM-based RAG systems, offering insights into achieving reliable and affordable RAG explainability.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04480
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Source Attribution in Retrieval-Augmented Generation
Nematov, Ikhtiyor
Kalai, Tarik
Kuzmenko, Elizaveta
Fugagnoli, Gabriele
Sacharidis, Dimitris
Hose, Katja
Sagi, Tomer
Machine Learning
Artificial Intelligence
While attribution methods, such as Shapley values, are widely used to explain the importance of features or training data in traditional machine learning, their application to Large Language Models (LLMs), particularly within Retrieval-Augmented Generation (RAG) systems, is nascent and challenging. The primary obstacle is the substantial computational cost, where each utility function evaluation involves an expensive LLM call, resulting in direct monetary and time expenses. This paper investigates the feasibility and effectiveness of adapting Shapley-based attribution to identify influential retrieved documents in RAG. We compare Shapley with more computationally tractable approximations and some existing attribution methods for LLM. Our work aims to: (1) systematically apply established attribution principles to the RAG document-level setting; (2) quantify how well SHAP approximations can mirror exact attributions while minimizing costly LLM interactions; and (3) evaluate their practical explainability in identifying critical documents, especially under complex inter-document relationships such as redundancy, complementarity, and synergy. This study seeks to bridge the gap between powerful attribution techniques and the practical constraints of LLM-based RAG systems, offering insights into achieving reliable and affordable RAG explainability.
title Source Attribution in Retrieval-Augmented Generation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.04480