HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Meng-Chieh, Zhu, Qi, Mavromatis, Costas, Han, Zhen, Adeshina, Soji, Ioannidis, Vassilis N., Rangwala, Huzefa, Faloutsos, Christos
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908390272270336
author Lee, Meng-Chieh
Zhu, Qi
Mavromatis, Costas
Han, Zhen
Adeshina, Soji
Ioannidis, Vassilis N.
Rangwala, Huzefa
Faloutsos, Christos
author_facet Lee, Meng-Chieh
Zhu, Qi
Mavromatis, Costas
Han, Zhen
Adeshina, Soji
Ioannidis, Vassilis N.
Rangwala, Huzefa
Faloutsos, Christos
contents Given a semi-structured knowledge base (SKB), where text documents are interconnected by relations, how can we effectively retrieve relevant information to answer user questions? Retrieval-Augmented Generation (RAG) retrieves documents to assist large language models (LLMs) in question answering; while Graph RAG (GRAG) uses structured knowledge bases as its knowledge source. However, many questions require both textual and relational information from SKB - referred to as "hybrid" questions - which complicates the retrieval process and underscores the need for a hybrid retrieval method that leverages both information. In this paper, through our empirical analysis, we identify key insights that show why existing methods may struggle with hybrid question answering (HQA) over SKB. Based on these insights, we propose HybGRAG for HQA consisting of a retriever bank and a critic module, with the following advantages: (1) Agentic, it automatically refines the output by incorporating feedback from the critic module, (2) Adaptive, it solves hybrid questions requiring both textual and relational information with the retriever bank, (3) Interpretable, it justifies decision making with intuitive refinement path, and (4) Effective, it surpasses all baselines on HQA benchmarks. In experiments on the STaRK benchmark, HybGRAG achieves significant performance gains, with an average relative improvement in Hit@1 of 51%.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16311
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases
Lee, Meng-Chieh
Zhu, Qi
Mavromatis, Costas
Han, Zhen
Adeshina, Soji
Ioannidis, Vassilis N.
Rangwala, Huzefa
Faloutsos, Christos
Machine Learning
Artificial Intelligence
Information Retrieval
Given a semi-structured knowledge base (SKB), where text documents are interconnected by relations, how can we effectively retrieve relevant information to answer user questions? Retrieval-Augmented Generation (RAG) retrieves documents to assist large language models (LLMs) in question answering; while Graph RAG (GRAG) uses structured knowledge bases as its knowledge source. However, many questions require both textual and relational information from SKB - referred to as "hybrid" questions - which complicates the retrieval process and underscores the need for a hybrid retrieval method that leverages both information. In this paper, through our empirical analysis, we identify key insights that show why existing methods may struggle with hybrid question answering (HQA) over SKB. Based on these insights, we propose HybGRAG for HQA consisting of a retriever bank and a critic module, with the following advantages: (1) Agentic, it automatically refines the output by incorporating feedback from the critic module, (2) Adaptive, it solves hybrid questions requiring both textual and relational information with the retriever bank, (3) Interpretable, it justifies decision making with intuitive refinement path, and (4) Effective, it surpasses all baselines on HQA benchmarks. In experiments on the STaRK benchmark, HybGRAG achieves significant performance gains, with an average relative improvement in Hit@1 of 51%.
title HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases
topic Machine Learning
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2412.16311