MSRS: Evaluating Multi-Source Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phanse, Rohan, Zhou, Yijie, Shi, Kejian, Zhang, Wencai, Liu, Yixin, Zhao, Yilun, Cohan, Arman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914011120926720
author Phanse, Rohan
Zhou, Yijie
Shi, Kejian
Zhang, Wencai
Liu, Yixin
Zhao, Yilun
Cohan, Arman
author_facet Phanse, Rohan
Zhou, Yijie
Shi, Kejian
Zhang, Wencai
Liu, Yixin
Zhao, Yilun
Cohan, Arman
contents Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or factoid-based. However, many real-world applications demand the ability to integrate and summarize information scattered across multiple sources, where no single source is sufficient to respond to the user's question. In such settings, the retrieval component of a RAG pipeline must recognize a variety of relevance signals, and the generation component must connect and synthesize information across multiple sources. We present a scalable framework for constructing evaluation benchmarks that challenge RAG systems to integrate information across distinct sources and generate long-form responses. Using our framework, we build two new benchmarks on Multi-Source Retrieval and Synthesis: MSRS-Story and MSRS-Meet, representing narrative synthesis and summarization tasks, respectively, that require retrieval from large collections. Our extensive experiments with various RAG pipelines -- including sparse and dense retrievers combined with frontier LLMs -- reveal that generation quality is highly dependent on retrieval effectiveness, which varies greatly by task. While multi-source synthesis proves challenging even in an oracle retrieval setting, we find that reasoning models significantly outperform standard LLMs at this distinct step.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
Phanse, Rohan
Zhou, Yijie
Shi, Kejian
Zhang, Wencai
Liu, Yixin
Zhao, Yilun
Cohan, Arman
Computation and Language
Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or factoid-based. However, many real-world applications demand the ability to integrate and summarize information scattered across multiple sources, where no single source is sufficient to respond to the user's question. In such settings, the retrieval component of a RAG pipeline must recognize a variety of relevance signals, and the generation component must connect and synthesize information across multiple sources. We present a scalable framework for constructing evaluation benchmarks that challenge RAG systems to integrate information across distinct sources and generate long-form responses. Using our framework, we build two new benchmarks on Multi-Source Retrieval and Synthesis: MSRS-Story and MSRS-Meet, representing narrative synthesis and summarization tasks, respectively, that require retrieval from large collections. Our extensive experiments with various RAG pipelines -- including sparse and dense retrievers combined with frontier LLMs -- reveal that generation quality is highly dependent on retrieval effectiveness, which varies greatly by task. While multi-source synthesis proves challenging even in an oracle retrieval setting, we find that reasoning models significantly outperform standard LLMs at this distinct step.
title MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
topic Computation and Language
url https://arxiv.org/abs/2508.20867