Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Bowen, Yoon, Jinsung, Han, Jiawei, Arik, Sercan O.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909342583750656
author Jin, Bowen
Yoon, Jinsung
Han, Jiawei
Arik, Sercan O.
author_facet Jin, Bowen
Yoon, Jinsung
Han, Jiawei
Arik, Sercan O.
contents Retrieval-augmented generation (RAG) empowers large language models (LLMs) to utilize external knowledge sources. The increasing capacity of LLMs to process longer input sequences opens up avenues for providing more retrieved information, to potentially enhance the quality of generated outputs. It is plausible to assume that a larger retrieval set would contain more relevant information (higher recall), that might result in improved performance. However, our empirical findings demonstrate that for many long-context LLMs, the quality of generated output initially improves first, but then subsequently declines as the number of retrieved passages increases. This paper investigates this phenomenon, identifying the detrimental impact of retrieved "hard negatives" as a key contributor. To mitigate this and enhance the robustness of long-context LLM-based RAG, we propose both training-free and training-based approaches. We first showcase the effectiveness of retrieval reordering as a simple yet powerful training-free optimization. Furthermore, we explore training-based methods, specifically RAG-specific implicit LLM fine-tuning and RAG-oriented fine-tuning with intermediate reasoning, demonstrating their capacity for substantial performance gains. Finally, we conduct a systematic analysis of design choices for these training-based methods, including data distribution, retriever selection, and training context length.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05983
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
Jin, Bowen
Yoon, Jinsung
Han, Jiawei
Arik, Sercan O.
Computation and Language
Artificial Intelligence
Machine Learning
Retrieval-augmented generation (RAG) empowers large language models (LLMs) to utilize external knowledge sources. The increasing capacity of LLMs to process longer input sequences opens up avenues for providing more retrieved information, to potentially enhance the quality of generated outputs. It is plausible to assume that a larger retrieval set would contain more relevant information (higher recall), that might result in improved performance. However, our empirical findings demonstrate that for many long-context LLMs, the quality of generated output initially improves first, but then subsequently declines as the number of retrieved passages increases. This paper investigates this phenomenon, identifying the detrimental impact of retrieved "hard negatives" as a key contributor. To mitigate this and enhance the robustness of long-context LLM-based RAG, we propose both training-free and training-based approaches. We first showcase the effectiveness of retrieval reordering as a simple yet powerful training-free optimization. Furthermore, we explore training-based methods, specifically RAG-specific implicit LLM fine-tuning and RAG-oriented fine-tuning with intermediate reasoning, demonstrating their capacity for substantial performance gains. Finally, we conduct a systematic analysis of design choices for these training-based methods, including data distribution, retriever selection, and training context length.
title Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.05983