Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Leung, Kin Kwan, Belbahri, Mouloud, Sui, Yi, Labach, Alex, Zhang, Xueying, Rose, Stephen Anthony, Cresswell, Jesse C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909990648807424
author Leung, Kin Kwan
Belbahri, Mouloud
Sui, Yi
Labach, Alex
Zhang, Xueying
Rose, Stephen Anthony
Cresswell, Jesse C.
author_facet Leung, Kin Kwan
Belbahri, Mouloud
Sui, Yi
Labach, Alex
Zhang, Xueying
Rose, Stephen Anthony
Cresswell, Jesse C.
contents Retrieval-augmented generation (RAG) is a prevalent approach for building LLM-based question-answering systems that can take advantage of external knowledge databases. Due to the complexity of real-world RAG systems, there are many potential causes for erroneous outputs. Understanding the range of errors that can occur in practice is crucial for robust deployment. We present a new taxonomy of the error types that can occur in realistic RAG systems, examples of each, and practical advice for addressing them. Additionally, we curate a dataset of erroneous RAG responses annotated by error types. We then propose an auto-evaluation method aligned with our taxonomy that can be used in practice to track and address errors during development. Code and data are available at https://github.com/layer6ai-labs/rag-error-classification.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
Leung, Kin Kwan
Belbahri, Mouloud
Sui, Yi
Labach, Alex
Zhang, Xueying
Rose, Stephen Anthony
Cresswell, Jesse C.
Computation and Language
Machine Learning
Retrieval-augmented generation (RAG) is a prevalent approach for building LLM-based question-answering systems that can take advantage of external knowledge databases. Due to the complexity of real-world RAG systems, there are many potential causes for erroneous outputs. Understanding the range of errors that can occur in practice is crucial for robust deployment. We present a new taxonomy of the error types that can occur in realistic RAG systems, examples of each, and practical advice for addressing them. Additionally, we curate a dataset of erroneous RAG responses annotated by error types. We then propose an auto-evaluation method aligned with our taxonomy that can be used in practice to track and address errors during development. Code and data are available at https://github.com/layer6ai-labs/rag-error-classification.
title Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.13975