Dynamic Context Selection for Retrieval-Augmented Generation: Mitigating Distractors and Positional Bias

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Iratni, Malika, Boughanem, Mohand, Dkaki, Taoufiq
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918249978920960
author Iratni, Malika
Boughanem, Mohand
Dkaki, Taoufiq
author_facet Iratni, Malika
Boughanem, Mohand
Dkaki, Taoufiq
contents Retrieval Augmented Generation (RAG) enhances language model performance by incorporating external knowledge retrieved from large corpora, which makes it highly suitable for tasks such as open domain question answering. Standard RAG systems typically rely on a fixed top k retrieval strategy, which can either miss relevant information or introduce semantically irrelevant passages, known as distractors, that degrade output quality. Additionally, the positioning of retrieved passages within the input context can influence the model attention and generation outcomes. Context placed in the middle tends to be overlooked, which is an issue known as the "lost in the middle" phenomenon. In this work, we systematically analyze the impact of distractors on generation quality, and quantify their effects under varying conditions. We also investigate how the position of relevant passages within the context window affects their influence on generation. Building on these insights, we propose a context-size classifier that dynamically predicts the optimal number of documents to retrieve based on query-specific informational needs. We integrate this approach into a full RAG pipeline, and demonstrate improved performance over fixed k baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14313
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Context Selection for Retrieval-Augmented Generation: Mitigating Distractors and Positional Bias
Iratni, Malika
Boughanem, Mohand
Dkaki, Taoufiq
Information Retrieval
Retrieval Augmented Generation (RAG) enhances language model performance by incorporating external knowledge retrieved from large corpora, which makes it highly suitable for tasks such as open domain question answering. Standard RAG systems typically rely on a fixed top k retrieval strategy, which can either miss relevant information or introduce semantically irrelevant passages, known as distractors, that degrade output quality. Additionally, the positioning of retrieved passages within the input context can influence the model attention and generation outcomes. Context placed in the middle tends to be overlooked, which is an issue known as the "lost in the middle" phenomenon. In this work, we systematically analyze the impact of distractors on generation quality, and quantify their effects under varying conditions. We also investigate how the position of relevant passages within the context window affects their influence on generation. Building on these insights, we propose a context-size classifier that dynamically predicts the optimal number of documents to retrieve based on query-specific informational needs. We integrate this approach into a full RAG pipeline, and demonstrate improved performance over fixed k baselines.
title Dynamic Context Selection for Retrieval-Augmented Generation: Mitigating Distractors and Positional Bias
topic Information Retrieval
url https://arxiv.org/abs/2512.14313