RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917751792074752 |
|---|---|
| author | Ru, Dongyu Qiu, Lin Hu, Xiangkun Zhang, Tianhang Shi, Peng Chang, Shuaichen Jiayang, Cheng Wang, Cunxiang Sun, Shichao Li, Huanyu Zhang, Zizhao Wang, Binjie Jiang, Jiarong He, Tong Wang, Zhiguo Liu, Pengfei Zhang, Yue Zhang, Zheng |
| author_facet | Ru, Dongyu Qiu, Lin Hu, Xiangkun Zhang, Tianhang Shi, Peng Chang, Shuaichen Jiayang, Cheng Wang, Cunxiang Sun, Shichao Li, Huanyu Zhang, Zizhao Wang, Binjie Jiang, Jiarong He, Tong Wang, Zhiguo Liu, Pengfei Zhang, Yue Zhang, Zheng |
| contents | Despite Retrieval-Augmented Generation (RAG) showing promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. This work has been open sourced at https://github.com/amazon-science/RAGChecker. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_08067 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation Ru, Dongyu Qiu, Lin Hu, Xiangkun Zhang, Tianhang Shi, Peng Chang, Shuaichen Jiayang, Cheng Wang, Cunxiang Sun, Shichao Li, Huanyu Zhang, Zizhao Wang, Binjie Jiang, Jiarong He, Tong Wang, Zhiguo Liu, Pengfei Zhang, Yue Zhang, Zheng Computation and Language Artificial Intelligence Despite Retrieval-Augmented Generation (RAG) showing promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. This work has been open sourced at https://github.com/amazon-science/RAGChecker. |
| title | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2408.08067 |