Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiatao, Hu, Xinyu, Yin, Xunjian, Wan, Xiaojun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929703186595840
author Li, Jiatao
Hu, Xinyu
Yin, Xunjian
Wan, Xiaojun
author_facet Li, Jiatao
Hu, Xinyu
Yin, Xunjian
Wan, Xiaojun
contents The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily focuses on optimizing the use of Self-Docs, with their inherent properties remaining underexplored. To bridge this gap, we first investigate the overall effectiveness of Self-Docs, identifying key factors that shape their contribution to RAG performance (RQ1). Building on these insights, we develop a taxonomy grounded in Systemic Functional Linguistics to compare the influence of various Self-Docs categories (RQ2) and explore strategies for combining them with external sources (RQ3). Our findings reveal which types of Self-Docs are most beneficial and offer practical guidelines for leveraging them to achieve significant improvements in knowledge-intensive question answering tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13192
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models
Li, Jiatao
Hu, Xinyu
Yin, Xunjian
Wan, Xiaojun
Computation and Language
The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily focuses on optimizing the use of Self-Docs, with their inherent properties remaining underexplored. To bridge this gap, we first investigate the overall effectiveness of Self-Docs, identifying key factors that shape their contribution to RAG performance (RQ1). Building on these insights, we develop a taxonomy grounded in Systemic Functional Linguistics to compare the influence of various Self-Docs categories (RQ2) and explore strategies for combining them with external sources (RQ3). Our findings reveal which types of Self-Docs are most beneficial and offer practical guidelines for leveraging them to achieve significant improvements in knowledge-intensive question answering tasks.
title Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2410.13192