Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909990648807424 |
|---|---|
| author | Leung, Kin Kwan Belbahri, Mouloud Sui, Yi Labach, Alex Zhang, Xueying Rose, Stephen Anthony Cresswell, Jesse C. |
| author_facet | Leung, Kin Kwan Belbahri, Mouloud Sui, Yi Labach, Alex Zhang, Xueying Rose, Stephen Anthony Cresswell, Jesse C. |
| contents | Retrieval-augmented generation (RAG) is a prevalent approach for building LLM-based question-answering systems that can take advantage of external knowledge databases. Due to the complexity of real-world RAG systems, there are many potential causes for erroneous outputs. Understanding the range of errors that can occur in practice is crucial for robust deployment. We present a new taxonomy of the error types that can occur in realistic RAG systems, examples of each, and practical advice for addressing them. Additionally, we curate a dataset of erroneous RAG responses annotated by error types. We then propose an auto-evaluation method aligned with our taxonomy that can be used in practice to track and address errors during development. Code and data are available at https://github.com/layer6ai-labs/rag-error-classification. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_13975 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems Leung, Kin Kwan Belbahri, Mouloud Sui, Yi Labach, Alex Zhang, Xueying Rose, Stephen Anthony Cresswell, Jesse C. Computation and Language Machine Learning Retrieval-augmented generation (RAG) is a prevalent approach for building LLM-based question-answering systems that can take advantage of external knowledge databases. Due to the complexity of real-world RAG systems, there are many potential causes for erroneous outputs. Understanding the range of errors that can occur in practice is crucial for robust deployment. We present a new taxonomy of the error types that can occur in realistic RAG systems, examples of each, and practical advice for addressing them. Additionally, we curate a dataset of erroneous RAG responses annotated by error types. We then propose an auto-evaluation method aligned with our taxonomy that can be used in practice to track and address errors during development. Code and data are available at https://github.com/layer6ai-labs/rag-error-classification. |
| title | Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.13975 |