A Reproducibility Study of LLM-Based Query Reformulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bigdeli, Amin, Rad, Radin Hamidi, Le, Hai Son, Incesu, Mert, Arabzadeh, Negar, Clarke, Charles L. A., Bagheri, Ebrahim
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911635343409152
author Bigdeli, Amin
Rad, Radin Hamidi
Le, Hai Son
Incesu, Mert
Arabzadeh, Negar
Clarke, Charles L. A.
Bagheri, Ebrahim
author_facet Bigdeli, Amin
Rad, Radin Hamidi
Le, Hai Son
Incesu, Mert
Arabzadeh, Negar
Clarke, Charles L. A.
Bagheri, Ebrahim
contents Large Language Models (LLMs) are now widely used for query reformulation and expansion in Information Retrieval, with many studies reporting substantial effectiveness gains. However, these results are typically obtained under heterogeneous experimental conditions, making it difficult to assess which findings are reproducible and which depend on specific implementation choices. In this work, we present a systematic reproducibility and comparative study of ten representative LLM-based query reformulation methods under a unified and strictly controlled experimental framework. We evaluate methods across two architectural LLM families at two parameter scales, three retrieval paradigms (lexical, learned sparse, and dense), and nine benchmark datasets spanning TREC Deep Learning and BEIR. Our results show that reformulation gains are strongly conditioned on the retrieval paradigm, that improvements observed under lexical retrieval do not consistently transfer to neural retrievers, and that larger LLMs do not uniformly yield better downstream performance. These findings clarify the stability and limits of reported gains in prior work. To enable transparent replication and ongoing comparison, we release all prompts, configurations, evaluation scripts, and run files through QueryGym, an open-source reformulation toolkit with a public leaderboard.\footnote{https://leaderboard.querygym.com}
format Preprint
id arxiv_https___arxiv_org_abs_2604_27421
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Reproducibility Study of LLM-Based Query Reformulation
Bigdeli, Amin
Rad, Radin Hamidi
Le, Hai Son
Incesu, Mert
Arabzadeh, Negar
Clarke, Charles L. A.
Bagheri, Ebrahim
Information Retrieval
Computation and Language
Large Language Models (LLMs) are now widely used for query reformulation and expansion in Information Retrieval, with many studies reporting substantial effectiveness gains. However, these results are typically obtained under heterogeneous experimental conditions, making it difficult to assess which findings are reproducible and which depend on specific implementation choices. In this work, we present a systematic reproducibility and comparative study of ten representative LLM-based query reformulation methods under a unified and strictly controlled experimental framework. We evaluate methods across two architectural LLM families at two parameter scales, three retrieval paradigms (lexical, learned sparse, and dense), and nine benchmark datasets spanning TREC Deep Learning and BEIR. Our results show that reformulation gains are strongly conditioned on the retrieval paradigm, that improvements observed under lexical retrieval do not consistently transfer to neural retrievers, and that larger LLMs do not uniformly yield better downstream performance. These findings clarify the stability and limits of reported gains in prior work. To enable transparent replication and ongoing comparison, we release all prompts, configurations, evaluation scripts, and run files through QueryGym, an open-source reformulation toolkit with a public leaderboard.\footnote{https://leaderboard.querygym.com}
title A Reproducibility Study of LLM-Based Query Reformulation
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2604.27421