Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929703186595840 |
|---|---|
| author | Li, Jiatao Hu, Xinyu Yin, Xunjian Wan, Xiaojun |
| author_facet | Li, Jiatao Hu, Xinyu Yin, Xunjian Wan, Xiaojun |
| contents | The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily focuses on optimizing the use of Self-Docs, with their inherent properties remaining underexplored. To bridge this gap, we first investigate the overall effectiveness of Self-Docs, identifying key factors that shape their contribution to RAG performance (RQ1). Building on these insights, we develop a taxonomy grounded in Systemic Functional Linguistics to compare the influence of various Self-Docs categories (RQ2) and explore strategies for combining them with external sources (RQ3). Our findings reveal which types of Self-Docs are most beneficial and offer practical guidelines for leveraging them to achieve significant improvements in knowledge-intensive question answering tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_13192 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models Li, Jiatao Hu, Xinyu Yin, Xunjian Wan, Xiaojun Computation and Language The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily focuses on optimizing the use of Self-Docs, with their inherent properties remaining underexplored. To bridge this gap, we first investigate the overall effectiveness of Self-Docs, identifying key factors that shape their contribution to RAG performance (RQ1). Building on these insights, we develop a taxonomy grounded in Systemic Functional Linguistics to compare the influence of various Self-Docs categories (RQ2) and explore strategies for combining them with external sources (RQ3). Our findings reveal which types of Self-Docs are most beneficial and offer practical guidelines for leveraging them to achieve significant improvements in knowledge-intensive question answering tasks. |
| title | Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2410.13192 |