Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gan, Aoran, Yu, Hao, Zhang, Kai, Liu, Qi, Yan, Wenyu, Huang, Zhenya, Tong, Shiwei, Hu, Guoping
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917992667807744
author Gan, Aoran
Yu, Hao
Zhang, Kai
Liu, Qi
Yan, Wenyu
Huang, Zhenya
Tong, Shiwei
Hu, Guoping
author_facet Gan, Aoran
Yu, Hao
Zhang, Kai
Liu, Qi
Yan, Wenyu
Huang, Zhenya
Tong, Shiwei
Hu, Guoping
contents Recent advancements in Retrieval-Augmented Generation (RAG) have revolutionized natural language processing by integrating Large Language Models (LLMs) with external information retrieval, enabling accurate, up-to-date, and verifiable text generation across diverse applications. However, evaluating RAG systems presents unique challenges due to their hybrid architecture that combines retrieval and generation components, as well as their dependence on dynamic knowledge sources in the LLM era. In response, this paper provides a comprehensive survey of RAG evaluation methods and frameworks, systematically reviewing traditional and emerging evaluation approaches, for system performance, factual accuracy, safety, and computational efficiency in the LLM era. We also compile and categorize the RAG-specific datasets and evaluation frameworks, conducting a meta-analysis of evaluation practices in high-impact RAG research. To the best of our knowledge, this work represents the most comprehensive survey for RAG evaluation, bridging traditional and LLM-driven methods, and serves as a critical resource for advancing RAG development.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
Gan, Aoran
Yu, Hao
Zhang, Kai
Liu, Qi
Yan, Wenyu
Huang, Zhenya
Tong, Shiwei
Hu, Guoping
Computation and Language
Recent advancements in Retrieval-Augmented Generation (RAG) have revolutionized natural language processing by integrating Large Language Models (LLMs) with external information retrieval, enabling accurate, up-to-date, and verifiable text generation across diverse applications. However, evaluating RAG systems presents unique challenges due to their hybrid architecture that combines retrieval and generation components, as well as their dependence on dynamic knowledge sources in the LLM era. In response, this paper provides a comprehensive survey of RAG evaluation methods and frameworks, systematically reviewing traditional and emerging evaluation approaches, for system performance, factual accuracy, safety, and computational efficiency in the LLM era. We also compile and categorize the RAG-specific datasets and evaluation frameworks, conducting a meta-analysis of evaluation practices in high-impact RAG research. To the best of our knowledge, this work represents the most comprehensive survey for RAG evaluation, bridging traditional and LLM-driven methods, and serves as a critical resource for advancing RAG development.
title Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
topic Computation and Language
url https://arxiv.org/abs/2504.14891