NeuCLIRBench: A Modern Evaluation Collection for Monolingual, Cross-Language, and Multilingual Information Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lawrie, Dawn, Mayfield, James, Yang, Eugene, Yates, Andrew, MacAvaney, Sean, Pradeep, Ronak, Miller, Scott, McNamee, Paul, Soldani, Luca
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911274776920064
author Lawrie, Dawn
Mayfield, James
Yang, Eugene
Yates, Andrew
MacAvaney, Sean
Pradeep, Ronak
Miller, Scott
McNamee, Paul
Soldani, Luca
author_facet Lawrie, Dawn
Mayfield, James
Yang, Eugene
Yates, Andrew
MacAvaney, Sean
Pradeep, Ronak
Miller, Scott
McNamee, Paul
Soldani, Luca
contents To measure advances in retrieval, test collections with relevance judgments that can faithfully distinguish systems are required. This paper presents NeuCLIRBench, an evaluation collection for cross-language and multilingual retrieval. The collection consists of documents written natively in Chinese, Persian, and Russian, as well as those same documents machine translated into English. The collection supports several retrieval scenarios including: monolingual retrieval in English, Chinese, Persian, or Russian; cross-language retrieval with English as the query language and one of the other three languages as the document language; and multilingual retrieval, again with English as the query language and relevant documents in all three languages. NeuCLIRBench combines the TREC NeuCLIR track topics of 2022, 2023, and 2024. The 250,128 judgments across approximately 150 queries for the monolingual and cross-language tasks and 100 queries for multilingual retrieval provide strong statistical discriminatory power to distinguish retrieval approaches. A fusion baseline of strong neural retrieval systems is included with the collection so that developers of reranking algorithms are no longer reliant on BM25 as their first-stage retriever. NeuCLIRBench is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NeuCLIRBench: A Modern Evaluation Collection for Monolingual, Cross-Language, and Multilingual Information Retrieval
Lawrie, Dawn
Mayfield, James
Yang, Eugene
Yates, Andrew
MacAvaney, Sean
Pradeep, Ronak
Miller, Scott
McNamee, Paul
Soldani, Luca
Information Retrieval
To measure advances in retrieval, test collections with relevance judgments that can faithfully distinguish systems are required. This paper presents NeuCLIRBench, an evaluation collection for cross-language and multilingual retrieval. The collection consists of documents written natively in Chinese, Persian, and Russian, as well as those same documents machine translated into English. The collection supports several retrieval scenarios including: monolingual retrieval in English, Chinese, Persian, or Russian; cross-language retrieval with English as the query language and one of the other three languages as the document language; and multilingual retrieval, again with English as the query language and relevant documents in all three languages. NeuCLIRBench combines the TREC NeuCLIR track topics of 2022, 2023, and 2024. The 250,128 judgments across approximately 150 queries for the monolingual and cross-language tasks and 100 queries for multilingual retrieval provide strong statistical discriminatory power to distinguish retrieval approaches. A fusion baseline of strong neural retrieval systems is included with the collection so that developers of reranking algorithms are no longer reliant on BM25 as their first-stage retriever. NeuCLIRBench is publicly available.
title NeuCLIRBench: A Modern Evaluation Collection for Monolingual, Cross-Language, and Multilingual Information Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2511.14758