EncouRAGe: Evaluating RAG Local, Fast, and Reliable

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Strich, Jan, Scharfenberg, Adeline, Biemann, Chris, Semmann, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914141977968640
author Strich, Jan
Scharfenberg, Adeline
Biemann, Chris
Semmann, Martin
author_facet Strich, Jan
Scharfenberg, Adeline
Biemann, Chris
Semmann, Martin
contents We introduce EncouRAGe, a comprehensive Python framework designed to streamline the development and evaluation of Retrieval-Augmented Generation (RAG) systems using Large Language Models (LLMs) and Embedding Models. EncouRAGe comprises five modular and extensible components: Type Manifest, RAG Factory, Inference, Vector Store, and Metrics, facilitating flexible experimentation and extensible development. The framework emphasizes scientific reproducibility, diverse evaluation metrics, and local deployment, enabling researchers to efficiently assess datasets within RAG workflows. This paper presents implementation details and an extensive evaluation across multiple benchmark datasets, including 25k QA pairs and over 51k documents. Our results show that RAG still underperforms compared to the Oracle Context, while Hybrid BM25 consistently achieves the best results across all four datasets. We further examine the effects of reranking, observing only marginal performance improvements accompanied by higher response latency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04696
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EncouRAGe: Evaluating RAG Local, Fast, and Reliable
Strich, Jan
Scharfenberg, Adeline
Biemann, Chris
Semmann, Martin
Computation and Language
Artificial Intelligence
Information Retrieval
We introduce EncouRAGe, a comprehensive Python framework designed to streamline the development and evaluation of Retrieval-Augmented Generation (RAG) systems using Large Language Models (LLMs) and Embedding Models. EncouRAGe comprises five modular and extensible components: Type Manifest, RAG Factory, Inference, Vector Store, and Metrics, facilitating flexible experimentation and extensible development. The framework emphasizes scientific reproducibility, diverse evaluation metrics, and local deployment, enabling researchers to efficiently assess datasets within RAG workflows. This paper presents implementation details and an extensive evaluation across multiple benchmark datasets, including 25k QA pairs and over 51k documents. Our results show that RAG still underperforms compared to the Oracle Context, while Hybrid BM25 consistently achieves the best results across all four datasets. We further examine the effects of reranking, observing only marginal performance improvements accompanied by higher response latency.
title EncouRAGe: Evaluating RAG Local, Fast, and Reliable
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2511.04696