How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Abdallah, Abdelrahman, Piryani, Bhawna, Mozafari, Jamshid, Ali, Mohammed, Jatowt, Adam
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918129541578752
author Abdallah, Abdelrahman
Piryani, Bhawna
Mozafari, Jamshid
Ali, Mohammed
Jatowt, Adam
author_facet Abdallah, Abdelrahman
Piryani, Bhawna
Mozafari, Jamshid
Ali, Mohammed
Jatowt, Adam
contents In this work, we present a systematic and comprehensive empirical evaluation of state-of-the-art reranking methods, encompassing large language model (LLM)-based, lightweight contextual, and zero-shot approaches, with respect to their performance in information retrieval tasks. We evaluate in total 22 methods, including 40 variants (depending on used LLM) across several established benchmarks, including TREC DL19, DL20, and BEIR, as well as a novel dataset designed to test queries unseen by pretrained models. Our primary goal is to determine, through controlled and fair comparisons, whether a performance disparity exists between LLM-based rerankers and their lightweight counterparts, particularly on novel queries, and to elucidate the underlying causes of any observed differences. To disentangle confounding factors, we analyze the effects of training data overlap, model architecture, and computational efficiency on reranking performance. Our findings indicate that while LLM-based rerankers demonstrate superior performance on familiar queries, their generalization ability to novel queries varies, with lightweight models offering comparable efficiency. We further identify that the novelty of queries significantly impacts reranking effectiveness, highlighting limitations in existing approaches. https://github.com/DataScienceUIBK/llm-reranking-generalization-study
format Preprint
id arxiv_https___arxiv_org_abs_2508_16757
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
Abdallah, Abdelrahman
Piryani, Bhawna
Mozafari, Jamshid
Ali, Mohammed
Jatowt, Adam
Computation and Language
Information Retrieval
In this work, we present a systematic and comprehensive empirical evaluation of state-of-the-art reranking methods, encompassing large language model (LLM)-based, lightweight contextual, and zero-shot approaches, with respect to their performance in information retrieval tasks. We evaluate in total 22 methods, including 40 variants (depending on used LLM) across several established benchmarks, including TREC DL19, DL20, and BEIR, as well as a novel dataset designed to test queries unseen by pretrained models. Our primary goal is to determine, through controlled and fair comparisons, whether a performance disparity exists between LLM-based rerankers and their lightweight counterparts, particularly on novel queries, and to elucidate the underlying causes of any observed differences. To disentangle confounding factors, we analyze the effects of training data overlap, model architecture, and computational efficiency on reranking performance. Our findings indicate that while LLM-based rerankers demonstrate superior performance on familiar queries, their generalization ability to novel queries varies, with lightweight models offering comparable efficiency. We further identify that the novelty of queries significantly impacts reranking effectiveness, highlighting limitations in existing approaches. https://github.com/DataScienceUIBK/llm-reranking-generalization-study
title How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2508.16757