Generative Query Expansion with Multilingual LLMs for Cross-Lingual Information Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Macmillan-Scott, Olivia, Goworek, Roksana, Özyiğit, Eda B.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915635327401984
author Macmillan-Scott, Olivia
Goworek, Roksana
Özyiğit, Eda B.
author_facet Macmillan-Scott, Olivia
Goworek, Roksana
Özyiğit, Eda B.
contents Query expansion is the reformulation of a user query by adding semantically related information, and is an essential component of monolingual and cross-lingual information retrieval used to ensure that relevant documents are not missed. Recently, multilingual large language models (mLLMs) have shifted query expansion from semantic augmentation with synonyms and related words to pseudo-document generation. Pseudo-documents both introduce additional relevant terms and bridge the gap between short queries and long documents, which is particularly beneficial in dense retrieval. This study evaluates recent mLLMs and fine-tuned variants across several generative expansion strategies to identify factors that drive cross-lingual retrieval performance. Results show that query length largely determines which prompting technique is effective, and that more elaborate prompts often do not yield further gains. Substantial linguistic disparities persist: cross-lingual query expansion can produce the largest improvements for languages with the weakest baselines, yet retrieval is especially poor between languages written in different scripts. Fine-tuning is found to lead to performance gains only when the training and test data are of similar format. These outcomes underline the need for more balanced multilingual and cross-lingual training and evaluation resources.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Query Expansion with Multilingual LLMs for Cross-Lingual Information Retrieval
Macmillan-Scott, Olivia
Goworek, Roksana
Özyiğit, Eda B.
Information Retrieval
Artificial Intelligence
Computation and Language
Query expansion is the reformulation of a user query by adding semantically related information, and is an essential component of monolingual and cross-lingual information retrieval used to ensure that relevant documents are not missed. Recently, multilingual large language models (mLLMs) have shifted query expansion from semantic augmentation with synonyms and related words to pseudo-document generation. Pseudo-documents both introduce additional relevant terms and bridge the gap between short queries and long documents, which is particularly beneficial in dense retrieval. This study evaluates recent mLLMs and fine-tuned variants across several generative expansion strategies to identify factors that drive cross-lingual retrieval performance. Results show that query length largely determines which prompting technique is effective, and that more elaborate prompts often do not yield further gains. Substantial linguistic disparities persist: cross-lingual query expansion can produce the largest improvements for languages with the weakest baselines, yet retrieval is especially poor between languages written in different scripts. Fine-tuning is found to lead to performance gains only when the training and test data are of similar format. These outcomes underline the need for more balanced multilingual and cross-lingual training and evaluation resources.
title Generative Query Expansion with Multilingual LLMs for Cross-Lingual Information Retrieval
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.19325