GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Yi, Zhao, Sendong, Li, Jianbo, Wang, Haochun, Qin, Bing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915303465680896
author Jiang, Yi
Zhao, Sendong
Li, Jianbo
Wang, Haochun
Qin, Bing
author_facet Jiang, Yi
Zhao, Sendong
Li, Jianbo
Wang, Haochun
Qin, Bing
contents The Retrieval-Augmented Generation (RAG) framework introduces a retrieval module to dynamically inject retrieved information into the input context of large language models (LLMs), and has demonstrated significant success in various NLP tasks. However, the current study points out that there is a preference gap between retrievers and LLMs in the RAG framework, which limit the further improvement of system performance. Some highly relevant passages may interfere with LLM reasoning because they contain complex or contradictory information; while some indirectly related or even inaccurate content may help LLM generate more accurate answers by providing suggestive information or logical clues. To solve this, we propose GainRAG, a novel approach that aligns the retriever's and LLM's preferences by defining a new metric, "gain", which measure how well an input passage contributes to correct outputs. Specifically, we propose a method to estimate these gain signals and train a middleware that aligns the preferences of the retriever and the LLM using only limited data. In addition, we introduce a pseudo-passage strategy to mitigate degradation. The experimental results on 6 datasets verify the effectiveness of GainRAG.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18710
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis
Jiang, Yi
Zhao, Sendong
Li, Jianbo
Wang, Haochun
Qin, Bing
Information Retrieval
Artificial Intelligence
The Retrieval-Augmented Generation (RAG) framework introduces a retrieval module to dynamically inject retrieved information into the input context of large language models (LLMs), and has demonstrated significant success in various NLP tasks. However, the current study points out that there is a preference gap between retrievers and LLMs in the RAG framework, which limit the further improvement of system performance. Some highly relevant passages may interfere with LLM reasoning because they contain complex or contradictory information; while some indirectly related or even inaccurate content may help LLM generate more accurate answers by providing suggestive information or logical clues. To solve this, we propose GainRAG, a novel approach that aligns the retriever's and LLM's preferences by defining a new metric, "gain", which measure how well an input passage contributes to correct outputs. Specifically, we propose a method to estimate these gain signals and train a middleware that aligns the preferences of the retriever and the LLM using only limited data. In addition, we introduce a pseudo-passage strategy to mitigate degradation. The experimental results on 6 datasets verify the effectiveness of GainRAG.
title GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2505.18710