Optimizing open-domain question answering with graph-based retrieval augmented generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cahoon, Joyce, Singh, Prerna, Litombe, Nick, Larson, Jonathan, Trinh, Ha, Zhu, Yiwen, Mueller, Andreas, Psallidas, Fotis, Curino, Carlo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929742687502336
author Cahoon, Joyce
Singh, Prerna
Litombe, Nick
Larson, Jonathan
Trinh, Ha
Zhu, Yiwen
Mueller, Andreas
Psallidas, Fotis
Curino, Carlo
author_facet Cahoon, Joyce
Singh, Prerna
Litombe, Nick
Larson, Jonathan
Trinh, Ha
Zhu, Yiwen
Mueller, Andreas
Psallidas, Fotis
Curino, Carlo
contents In this work, we benchmark various graph-based retrieval-augmented generation (RAG) systems across a broad spectrum of query types, including OLTP-style (fact-based) and OLAP-style (thematic) queries, to address the complex demands of open-domain question answering (QA). Traditional RAG methods often fall short in handling nuanced, multi-document synthesis tasks. By structuring knowledge as graphs, we can facilitate the retrieval of context that captures greater semantic depth and enhances language model operations. We explore graph-based RAG methodologies and introduce TREX, a novel, cost-effective alternative that combines graph-based and vector-based retrieval techniques. Our benchmarking across four diverse datasets highlights the strengths of different RAG methodologies, demonstrates TREX's ability to handle multiple open-domain QA types, and reveals the limitations of current evaluation methods. In a real-world technical support case study, we demonstrate how TREX solutions can surpass conventional vector-based RAG in efficiently synthesizing data from heterogeneous sources. Our findings underscore the potential of augmenting large language models with advanced retrieval and orchestration capabilities, advancing scalable, graph-based AI solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing open-domain question answering with graph-based retrieval augmented generation
Cahoon, Joyce
Singh, Prerna
Litombe, Nick
Larson, Jonathan
Trinh, Ha
Zhu, Yiwen
Mueller, Andreas
Psallidas, Fotis
Curino, Carlo
Information Retrieval
H.3.3; I.2.7
In this work, we benchmark various graph-based retrieval-augmented generation (RAG) systems across a broad spectrum of query types, including OLTP-style (fact-based) and OLAP-style (thematic) queries, to address the complex demands of open-domain question answering (QA). Traditional RAG methods often fall short in handling nuanced, multi-document synthesis tasks. By structuring knowledge as graphs, we can facilitate the retrieval of context that captures greater semantic depth and enhances language model operations. We explore graph-based RAG methodologies and introduce TREX, a novel, cost-effective alternative that combines graph-based and vector-based retrieval techniques. Our benchmarking across four diverse datasets highlights the strengths of different RAG methodologies, demonstrates TREX's ability to handle multiple open-domain QA types, and reveals the limitations of current evaluation methods. In a real-world technical support case study, we demonstrate how TREX solutions can surpass conventional vector-based RAG in efficiently synthesizing data from heterogeneous sources. Our findings underscore the potential of augmenting large language models with advanced retrieval and orchestration capabilities, advancing scalable, graph-based AI solutions.
title Optimizing open-domain question answering with graph-based retrieval augmented generation
topic Information Retrieval
H.3.3; I.2.7
url https://arxiv.org/abs/2503.02922