Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alaofi, Marwah, Gallagher, Luke, Sanderson, Mark, Scholer, Falk, Thomas, Paul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915129208078336
author Alaofi, Marwah
Gallagher, Luke
Sanderson, Mark
Scholer, Falk
Thomas, Paul
author_facet Alaofi, Marwah
Gallagher, Luke
Sanderson, Mark
Scholer, Falk
Thomas, Paul
contents This paper explores the utility of a Large Language Model (LLM) to automatically generate queries and query variants from a description of an information need. Given a set of information needs described as backstories, we explore how similar the queries generated by the LLM are to those generated by humans. We quantify the similarity using different metrics and examine how the use of each set would contribute to document pooling when building test collections. Our results show potential in using LLMs to generate query variants. While they may not fully capture the wide variety of human-generated variants, they generate similar sets of relevant documents, reaching up to 71.1% overlap at a pool depth of 100.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17981
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
Alaofi, Marwah
Gallagher, Luke
Sanderson, Mark
Scholer, Falk
Thomas, Paul
Information Retrieval
This paper explores the utility of a Large Language Model (LLM) to automatically generate queries and query variants from a description of an information need. Given a set of information needs described as backstories, we explore how similar the queries generated by the LLM are to those generated by humans. We quantify the similarity using different metrics and examine how the use of each set would contribute to document pooling when building test collections. Our results show potential in using LLMs to generate query variants. While they may not fully capture the wide variety of human-generated variants, they generate similar sets of relevant documents, reaching up to 71.1% overlap at a pool depth of 100.
title Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
topic Information Retrieval
url https://arxiv.org/abs/2501.17981