Optimizing open-domain question answering with graph-based retrieval augmented generation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929742687502336 |
|---|---|
| author | Cahoon, Joyce Singh, Prerna Litombe, Nick Larson, Jonathan Trinh, Ha Zhu, Yiwen Mueller, Andreas Psallidas, Fotis Curino, Carlo |
| author_facet | Cahoon, Joyce Singh, Prerna Litombe, Nick Larson, Jonathan Trinh, Ha Zhu, Yiwen Mueller, Andreas Psallidas, Fotis Curino, Carlo |
| contents | In this work, we benchmark various graph-based retrieval-augmented generation (RAG) systems across a broad spectrum of query types, including OLTP-style (fact-based) and OLAP-style (thematic) queries, to address the complex demands of open-domain question answering (QA). Traditional RAG methods often fall short in handling nuanced, multi-document synthesis tasks. By structuring knowledge as graphs, we can facilitate the retrieval of context that captures greater semantic depth and enhances language model operations. We explore graph-based RAG methodologies and introduce TREX, a novel, cost-effective alternative that combines graph-based and vector-based retrieval techniques. Our benchmarking across four diverse datasets highlights the strengths of different RAG methodologies, demonstrates TREX's ability to handle multiple open-domain QA types, and reveals the limitations of current evaluation methods.
In a real-world technical support case study, we demonstrate how TREX solutions can surpass conventional vector-based RAG in efficiently synthesizing data from heterogeneous sources. Our findings underscore the potential of augmenting large language models with advanced retrieval and orchestration capabilities, advancing scalable, graph-based AI solutions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_02922 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Optimizing open-domain question answering with graph-based retrieval augmented generation Cahoon, Joyce Singh, Prerna Litombe, Nick Larson, Jonathan Trinh, Ha Zhu, Yiwen Mueller, Andreas Psallidas, Fotis Curino, Carlo Information Retrieval H.3.3; I.2.7 In this work, we benchmark various graph-based retrieval-augmented generation (RAG) systems across a broad spectrum of query types, including OLTP-style (fact-based) and OLAP-style (thematic) queries, to address the complex demands of open-domain question answering (QA). Traditional RAG methods often fall short in handling nuanced, multi-document synthesis tasks. By structuring knowledge as graphs, we can facilitate the retrieval of context that captures greater semantic depth and enhances language model operations. We explore graph-based RAG methodologies and introduce TREX, a novel, cost-effective alternative that combines graph-based and vector-based retrieval techniques. Our benchmarking across four diverse datasets highlights the strengths of different RAG methodologies, demonstrates TREX's ability to handle multiple open-domain QA types, and reveals the limitations of current evaluation methods. In a real-world technical support case study, we demonstrate how TREX solutions can surpass conventional vector-based RAG in efficiently synthesizing data from heterogeneous sources. Our findings underscore the potential of augmenting large language models with advanced retrieval and orchestration capabilities, advancing scalable, graph-based AI solutions. |
| title | Optimizing open-domain question answering with graph-based retrieval augmented generation |
| topic | Information Retrieval H.3.3; I.2.7 |
| url | https://arxiv.org/abs/2503.02922 |