Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kuissi, Nathan, Subrahmanyan, Suraj, Thakur, Nandan, Lin, Jimmy
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917315183902720
author Kuissi, Nathan
Subrahmanyan, Suraj
Thakur, Nandan
Lin, Jimmy
author_facet Kuissi, Nathan
Subrahmanyan, Suraj
Thakur, Nandan
Lin, Jimmy
contents Information retrieval (IR) benchmarks typically follow the Cranfield paradigm, relying on static and predefined corpora. However, temporal changes in technical corpora, such as API deprecations and code reorganizations, can render existing benchmarks stale. In our work, we investigate how temporal corpus drift affects FreshStack, a retrieval benchmark focused on technical domains. We examine two independent corpus snapshots of FreshStack from October 2024 and October 2025 to answer questions about LangChain. Our analysis shows that all but one query posed in 2024 remain fully supported by the 2025 corpus, as relevant documents "migrate" from LangChain to competitor repositories, such as LlamaIndex. Next, we compare the accuracy of retrieval models on both snapshots and observe only minor shifts in model rankings, with overall strong correlation of up to 0.978 Kendall $τ$ at Recall@50. These results suggest that retrieval benchmarks re-judged with evolving temporal corpora can remain reliable for retrieval evaluation. We publicly release all our artifacts at https://github.com/fresh-stack/driftbench.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04532
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
Kuissi, Nathan
Subrahmanyan, Suraj
Thakur, Nandan
Lin, Jimmy
Information Retrieval
Artificial Intelligence
Computation and Language
Information retrieval (IR) benchmarks typically follow the Cranfield paradigm, relying on static and predefined corpora. However, temporal changes in technical corpora, such as API deprecations and code reorganizations, can render existing benchmarks stale. In our work, we investigate how temporal corpus drift affects FreshStack, a retrieval benchmark focused on technical domains. We examine two independent corpus snapshots of FreshStack from October 2024 and October 2025 to answer questions about LangChain. Our analysis shows that all but one query posed in 2024 remain fully supported by the 2025 corpus, as relevant documents "migrate" from LangChain to competitor repositories, such as LlamaIndex. Next, we compare the accuracy of retrieval models on both snapshots and observe only minor shifts in model rankings, with overall strong correlation of up to 0.978 Kendall $τ$ at Recall@50. These results suggest that retrieval benchmarks re-judged with evolving temporal corpora can remain reliable for retrieval evaluation. We publicly release all our artifacts at https://github.com/fresh-stack/driftbench.
title Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.04532