Do RAG Systems Really Suffer From Positional Bias?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cuconasu, Florin, Filice, Simone, Horowitz, Guy, Maarek, Yoelle, Silvestri, Fabrizio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915537714413568
author Cuconasu, Florin
Filice, Simone
Horowitz, Guy
Maarek, Yoelle
Silvestri, Fabrizio
author_facet Cuconasu, Florin
Filice, Simone
Horowitz, Guy
Maarek, Yoelle
Silvestri, Fabrizio
contents Retrieval Augmented Generation enhances LLM accuracy by adding passages retrieved from an external corpus to the LLM prompt. This paper investigates how positional bias - the tendency of LLMs to weight information differently based on its position in the prompt - affects not only the LLM's capability to capitalize on relevant passages, but also its susceptibility to distracting passages. Through extensive experiments on three benchmarks, we show how state-of-the-art retrieval pipelines, while attempting to retrieve relevant passages, systematically bring highly distracting ones to the top ranks, with over 60% of queries containing at least one highly distracting passage among the top-10 retrieved passages. As a result, the impact of the LLM positional bias, which in controlled settings is often reported as very prominent by related works, is actually marginal in real scenarios since both relevant and distracting passages are, in turn, penalized. Indeed, our findings reveal that sophisticated strategies that attempt to rearrange the passages based on LLM positional preferences do not perform better than random shuffling.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15561
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do RAG Systems Really Suffer From Positional Bias?
Cuconasu, Florin
Filice, Simone
Horowitz, Guy
Maarek, Yoelle
Silvestri, Fabrizio
Computation and Language
Information Retrieval
Retrieval Augmented Generation enhances LLM accuracy by adding passages retrieved from an external corpus to the LLM prompt. This paper investigates how positional bias - the tendency of LLMs to weight information differently based on its position in the prompt - affects not only the LLM's capability to capitalize on relevant passages, but also its susceptibility to distracting passages. Through extensive experiments on three benchmarks, we show how state-of-the-art retrieval pipelines, while attempting to retrieve relevant passages, systematically bring highly distracting ones to the top ranks, with over 60% of queries containing at least one highly distracting passage among the top-10 retrieved passages. As a result, the impact of the LLM positional bias, which in controlled settings is often reported as very prominent by related works, is actually marginal in real scenarios since both relevant and distracting passages are, in turn, penalized. Indeed, our findings reveal that sophisticated strategies that attempt to rearrange the passages based on LLM positional preferences do not perform better than random shuffling.
title Do RAG Systems Really Suffer From Positional Bias?
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2505.15561