Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fang, Hanpei, Tao, Sijie, Chen, Nuo, Chang, Kai-Xin, Sakai, Tetsuya
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916949183692800
author Fang, Hanpei
Tao, Sijie
Chen, Nuo
Chang, Kai-Xin
Sakai, Tetsuya
author_facet Fang, Hanpei
Tao, Sijie
Chen, Nuo
Chang, Kai-Xin
Sakai, Tetsuya
contents Large language models (LLMs) are increasingly deployed in information systems, including being used as second-stage rerankers in information retrieval pipelines, yet their susceptibility to recency bias has received little attention. We investigate whether LLMs implicitly favour newer documents by prepending artificial publication dates to passages in the TREC Deep Learning passage retrieval collections in 2021 (DL21) and 2022 (DL22). Across seven models, GPT-3.5-turbo, GPT-4o, GPT-4, LLaMA-3 8B/70B, and Qwen-2.5 7B/72B, "fresh" passages are consistently promoted, shifting the Top-10's mean publication year forward by up to 4.78 years and moving individual items by as many as 95 ranks in our listwise reranking experiments. Although larger models attenuate the effect, none eliminate it. We also observe that the preference of LLMs between two passages with an identical relevance level can be reversed by up to 25% on average after date injection in our pairwise preference experiments. These findings provide quantitative evidence of a pervasive recency bias in LLMs and highlight the importance of effective bias-mitigation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11353
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking
Fang, Hanpei
Tao, Sijie
Chen, Nuo
Chang, Kai-Xin
Sakai, Tetsuya
Information Retrieval
Large language models (LLMs) are increasingly deployed in information systems, including being used as second-stage rerankers in information retrieval pipelines, yet their susceptibility to recency bias has received little attention. We investigate whether LLMs implicitly favour newer documents by prepending artificial publication dates to passages in the TREC Deep Learning passage retrieval collections in 2021 (DL21) and 2022 (DL22). Across seven models, GPT-3.5-turbo, GPT-4o, GPT-4, LLaMA-3 8B/70B, and Qwen-2.5 7B/72B, "fresh" passages are consistently promoted, shifting the Top-10's mean publication year forward by up to 4.78 years and moving individual items by as many as 95 ranks in our listwise reranking experiments. Although larger models attenuate the effect, none eliminate it. We also observe that the preference of LLMs between two passages with an identical relevance level can be reversed by up to 25% on average after date injection in our pairwise preference experiments. These findings provide quantitative evidence of a pervasive recency bias in LLMs and highlight the importance of effective bias-mitigation strategies.
title Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking
topic Information Retrieval
url https://arxiv.org/abs/2509.11353