Evaluating Large Language Models for Cross-Lingual Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Longfei, Hong, Pingjun, Kraus, Oliver, Plank, Barbara, Litschko, Robert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908545479344128
author Zuo, Longfei
Hong, Pingjun
Kraus, Oliver
Plank, Barbara
Litschko, Robert
author_facet Zuo, Longfei
Hong, Pingjun
Kraus, Oliver
Plank, Barbara
Litschko, Robert
contents Multi-stage information retrieval (IR) has become a widely-adopted paradigm in search. While Large Language Models (LLMs) have been extensively evaluated as second-stage reranking models for monolingual IR, a systematic large-scale comparison is still lacking for cross-lingual IR (CLIR). Moreover, while prior work shows that LLM-based rerankers improve CLIR performance, their evaluation setup relies on lexical retrieval with machine translation (MT) for the first stage. This is not only prohibitively expensive but also prone to error propagation across stages. Our evaluation on passage-level and document-level CLIR reveals that further gains can be achieved with multilingual bi-encoders as first-stage retrievers and that the benefits of translation diminishes with stronger reranking models. We further show that pairwise rerankers based on instruction-tuned LLMs perform competitively with listwise rerankers. To the best of our knowledge, we are the first to study the interaction between retrievers and rerankers in two-stage CLIR with LLMs. Our findings reveal that, without MT, current state-of-the-art rerankers fall severely short when directly applied in CLIR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Large Language Models for Cross-Lingual Retrieval
Zuo, Longfei
Hong, Pingjun
Kraus, Oliver
Plank, Barbara
Litschko, Robert
Computation and Language
Information Retrieval
Multi-stage information retrieval (IR) has become a widely-adopted paradigm in search. While Large Language Models (LLMs) have been extensively evaluated as second-stage reranking models for monolingual IR, a systematic large-scale comparison is still lacking for cross-lingual IR (CLIR). Moreover, while prior work shows that LLM-based rerankers improve CLIR performance, their evaluation setup relies on lexical retrieval with machine translation (MT) for the first stage. This is not only prohibitively expensive but also prone to error propagation across stages. Our evaluation on passage-level and document-level CLIR reveals that further gains can be achieved with multilingual bi-encoders as first-stage retrievers and that the benefits of translation diminishes with stronger reranking models. We further show that pairwise rerankers based on instruction-tuned LLMs perform competitively with listwise rerankers. To the best of our knowledge, we are the first to study the interaction between retrievers and rerankers in two-stage CLIR with LLMs. Our findings reveal that, without MT, current state-of-the-art rerankers fall severely short when directly applied in CLIR.
title Evaluating Large Language Models for Cross-Lingual Retrieval
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2509.14749