Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hamman, Faisal, Zhu, Chenyang, Kumar, Anoop, Peng, Xujun, Dutta, Sanghamitra, Liu, Daben, Samuel, Alfy
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915533909131264
author Hamman, Faisal
Zhu, Chenyang
Kumar, Anoop
Peng, Xujun
Dutta, Sanghamitra
Liu, Daben
Samuel, Alfy
author_facet Hamman, Faisal
Zhu, Chenyang
Kumar, Anoop
Peng, Xujun
Dutta, Sanghamitra
Liu, Daben
Samuel, Alfy
contents RAG systems are increasingly deployed in high-stakes domains where users expect outputs to be consistent across semantically equivalent queries. However, existing systems often exhibit significant inconsistencies due to variability in both the retriever and generator (LLM), undermining trust and reliability. In this work, we focus on information consistency, i.e., the requirement that outputs convey the same core content across semantically equivalent inputs. We introduce a principled evaluation framework that decomposes RAG consistency into retriever-level, generator-level, and end-to-end components, helping identify inconsistency sources. To improve consistency, we propose Paraphrased Set Group Relative Policy Optimization (PS-GRPO), an RL approach that leverages multiple rollouts across paraphrased set to assign group similarity rewards. We leverage PS-GRPO to achieve Information Consistent RAG (Con-RAG), training the generator to produce consistent outputs across paraphrased queries and remain robust to retrieval-induced variability. Because exact reward computation over paraphrase sets is computationally expensive, we also introduce a scalable approximation method that retains effectiveness while enabling efficient, large-scale training. Empirical evaluations across short-form, multi-hop, and long-form QA benchmarks demonstrate that Con-RAG significantly improves both consistency and accuracy over strong baselines, even in the absence of explicit ground-truth supervision. Our work provides practical solutions for evaluating and building reliable RAG systems for safety-critical deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
Hamman, Faisal
Zhu, Chenyang
Kumar, Anoop
Peng, Xujun
Dutta, Sanghamitra
Liu, Daben
Samuel, Alfy
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
RAG systems are increasingly deployed in high-stakes domains where users expect outputs to be consistent across semantically equivalent queries. However, existing systems often exhibit significant inconsistencies due to variability in both the retriever and generator (LLM), undermining trust and reliability. In this work, we focus on information consistency, i.e., the requirement that outputs convey the same core content across semantically equivalent inputs. We introduce a principled evaluation framework that decomposes RAG consistency into retriever-level, generator-level, and end-to-end components, helping identify inconsistency sources. To improve consistency, we propose Paraphrased Set Group Relative Policy Optimization (PS-GRPO), an RL approach that leverages multiple rollouts across paraphrased set to assign group similarity rewards. We leverage PS-GRPO to achieve Information Consistent RAG (Con-RAG), training the generator to produce consistent outputs across paraphrased queries and remain robust to retrieval-induced variability. Because exact reward computation over paraphrase sets is computationally expensive, we also introduce a scalable approximation method that retains effectiveness while enabling efficient, large-scale training. Empirical evaluations across short-form, multi-hop, and long-form QA benchmarks demonstrate that Con-RAG significantly improves both consistency and accuracy over strong baselines, even in the absence of explicit ground-truth supervision. Our work provides practical solutions for evaluating and building reliable RAG systems for safety-critical deployments.
title Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2510.04392