Constructing and Evaluating Declarative RAG Pipelines in PyTerrier

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Macdonald, Craig, Fang, Jinyuan, Parry, Andrew, Meng, Zaiqiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911002212171776
author Macdonald, Craig
Fang, Jinyuan
Parry, Andrew
Meng, Zaiqiao
author_facet Macdonald, Craig
Fang, Jinyuan
Parry, Andrew
Meng, Zaiqiao
contents Search engines often follow a pipeline architecture, where complex but effective reranking components are used to refine the results of an initial retrieval. Retrieval augmented generation (RAG) is an exciting application of the pipeline architecture, where the final component generates a coherent answer for the users from the retrieved documents. In this demo paper, we describe how such RAG pipelines can be formulated in the declarative PyTerrier architecture, and the advantages of doing so. Our PyTerrier-RAG extension for PyTerrier provides easy access to standard RAG datasets and evaluation measures, state-of-the-art LLM readers, and using PyTerrier's unique operator notation, easy-to-build pipelines. We demonstrate the succinctness of indexing and RAG pipelines on standard datasets (including Natural Questions) and how to build on the larger PyTerrier ecosystem with state-of-the-art sparse, learned-sparse, and dense retrievers, and other neural rankers.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10802
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
Macdonald, Craig
Fang, Jinyuan
Parry, Andrew
Meng, Zaiqiao
Information Retrieval
Search engines often follow a pipeline architecture, where complex but effective reranking components are used to refine the results of an initial retrieval. Retrieval augmented generation (RAG) is an exciting application of the pipeline architecture, where the final component generates a coherent answer for the users from the retrieved documents. In this demo paper, we describe how such RAG pipelines can be formulated in the declarative PyTerrier architecture, and the advantages of doing so. Our PyTerrier-RAG extension for PyTerrier provides easy access to standard RAG datasets and evaluation measures, state-of-the-art LLM readers, and using PyTerrier's unique operator notation, easy-to-build pipelines. We demonstrate the succinctness of indexing and RAG pipelines on standard datasets (including Natural Questions) and how to build on the larger PyTerrier ecosystem with state-of-the-art sparse, learned-sparse, and dense retrievers, and other neural rankers.
title Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
topic Information Retrieval
url https://arxiv.org/abs/2506.10802