On the impact of retrieved content representations in RAG Pipelines

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ross, Jonathan J, Koopman, Bevan, van der Vegt, Anton, Zuccon, Guido
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917547079630848
author Ross, Jonathan J
Koopman, Bevan
van der Vegt, Anton
Zuccon, Guido
author_facet Ross, Jonathan J
Koopman, Bevan
van der Vegt, Anton
Zuccon, Guido
contents Retrieval-Augmented Generation (RAG) supplements a language model's input with retrieved documents, yet most RAG pipelines inherit retrieval components designed for human readers. How retrieved content should be represented when the consumer is a large language model (LLM) rather than a human is less well understood. Recent work has proposed transformations of retrieved content and identified properties that affect generation, but each examines a single transformation or property in isolation, leaving open which features of a document's representation matter most. We address this with a controlled comparison: holding retrieval fixed, we vary only the representation of retrieved documents, comparing an original baseline against thirteen transformations spanning selection, summarisation, and reformulation, in query-dependent and query-independent variants. Across these fourteen representations we measure question-answering accuracy for four generators, and for each representation we also measure answer retention: whether a known answer-bearing document still supports its answer after transformation. We find that answer retention is the primary determinant of generator accuracy; notably, when retention is high, a representation's wording, structure, length, and query-dependence have limited effect. This suggests that accuracy gains attributed to specific mechanisms in prior work may be partly explained by how well those mechanisms preserve answer-bearing content, an attribution that cannot be settled without controlling for retention.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30790
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the impact of retrieved content representations in RAG Pipelines
Ross, Jonathan J
Koopman, Bevan
van der Vegt, Anton
Zuccon, Guido
Information Retrieval
Artificial Intelligence
Computation and Language
Retrieval-Augmented Generation (RAG) supplements a language model's input with retrieved documents, yet most RAG pipelines inherit retrieval components designed for human readers. How retrieved content should be represented when the consumer is a large language model (LLM) rather than a human is less well understood. Recent work has proposed transformations of retrieved content and identified properties that affect generation, but each examines a single transformation or property in isolation, leaving open which features of a document's representation matter most. We address this with a controlled comparison: holding retrieval fixed, we vary only the representation of retrieved documents, comparing an original baseline against thirteen transformations spanning selection, summarisation, and reformulation, in query-dependent and query-independent variants. Across these fourteen representations we measure question-answering accuracy for four generators, and for each representation we also measure answer retention: whether a known answer-bearing document still supports its answer after transformation. We find that answer retention is the primary determinant of generator accuracy; notably, when retention is high, a representation's wording, structure, length, and query-dependence have limited effect. This suggests that accuracy gains attributed to specific mechanisms in prior work may be partly explained by how well those mechanisms preserve answer-bearing content, an attribution that cannot be settled without controlling for retention.
title On the impact of retrieved content representations in RAG Pipelines
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.30790