Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shi, Zhengliang, Sun, Weiwei, Gao, Shen, Ren, Pengjie, Chen, Zhumin, Ren, Zhaochun
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912029897392128
author Shi, Zhengliang
Sun, Weiwei
Gao, Shen
Ren, Pengjie
Chen, Zhumin
Ren, Zhaochun
author_facet Shi, Zhengliang
Sun, Weiwei
Gao, Shen
Ren, Pengjie
Chen, Zhumin
Ren, Zhaochun
contents Multi-Hop Question Answering (MHQA) tasks present a significant challenge for large language models (LLMs) due to the intensive knowledge required. Current solutions, like Retrieval-Augmented Generation, typically retrieve potential documents from an external corpus to read an answer. However, the performance of this retrieve-then-read paradigm is constrained by the retriever and the inevitable noise in the retrieved documents. To mitigate these challenges, we introduce a novel generate-then-ground (GenGround) framework, synergizing the parametric knowledge of LLMs and external documents to solve a multi-hop question. GenGround empowers LLMs to alternate two phases until the final answer is derived: (1) formulate a simpler, single-hop question and directly generate the answer; (2) ground the question-answer pair in retrieved documents, amending any wrong predictions in the answer. We also propose an instructional grounding distillation method to generalize our method into smaller models. Extensive experiments conducted on four datasets illustrate the superiority of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering
Shi, Zhengliang
Sun, Weiwei
Gao, Shen
Ren, Pengjie
Chen, Zhumin
Ren, Zhaochun
Computation and Language
Information Retrieval
Multi-Hop Question Answering (MHQA) tasks present a significant challenge for large language models (LLMs) due to the intensive knowledge required. Current solutions, like Retrieval-Augmented Generation, typically retrieve potential documents from an external corpus to read an answer. However, the performance of this retrieve-then-read paradigm is constrained by the retriever and the inevitable noise in the retrieved documents. To mitigate these challenges, we introduce a novel generate-then-ground (GenGround) framework, synergizing the parametric knowledge of LLMs and external documents to solve a multi-hop question. GenGround empowers LLMs to alternate two phases until the final answer is derived: (1) formulate a simpler, single-hop question and directly generate the answer; (2) ground the question-answer pair in retrieved documents, amending any wrong predictions in the answer. We also propose an instructional grounding distillation method to generalize our method into smaller models. Extensive experiments conducted on four datasets illustrate the superiority of our method.
title Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2406.14891