BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wojtasik, Konrad, Shishkin, Vadim, Wołowiec, Kacper, Janz, Arkadiusz, Piasecki, Maciej
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917667308306432
author Wojtasik, Konrad
Shishkin, Vadim
Wołowiec, Kacper
Janz, Arkadiusz
Piasecki, Maciej
author_facet Wojtasik, Konrad
Shishkin, Vadim
Wołowiec, Kacper
Janz, Arkadiusz
Piasecki, Maciej
contents The BEIR dataset is a large, heterogeneous benchmark for Information Retrieval (IR) in zero-shot settings, garnering considerable attention within the research community. However, BEIR and analogous datasets are predominantly restricted to the English language. Our objective is to establish extensive large-scale resources for IR in the Polish language, thereby advancing the research in this NLP area. In this work, inspired by mMARCO and Mr.~TyDi datasets, we translated all accessible open IR datasets into Polish, and we introduced the BEIR-PL benchmark -- a new benchmark which comprises 13 datasets, facilitating further development, training and evaluation of modern Polish language models for IR tasks. We executed an evaluation and comparison of numerous IR models on the newly introduced BEIR-PL benchmark. Furthermore, we publish pre-trained open IR models for Polish language,d marking a pioneering development in this field. Additionally, the evaluation revealed that BM25 achieved significantly lower scores for Polish than for English, which can be attributed to high inflection and intricate morphological structure of the Polish language. Finally, we trained various re-ranking models to enhance the BM25 retrieval, and we compared their performance to identify their unique characteristic features. To ensure accurate model comparisons, it is necessary to scrutinise individual results rather than to average across the entire benchmark. Thus, we thoroughly analysed the outcomes of IR models in relation to each individual data subset encompassed by the BEIR benchmark. The benchmark data is available at URL {\bf https://huggingface.co/clarin-knext}.
format Preprint
id arxiv_https___arxiv_org_abs_2305_19840
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
Wojtasik, Konrad
Shishkin, Vadim
Wołowiec, Kacper
Janz, Arkadiusz
Piasecki, Maciej
Information Retrieval
Artificial Intelligence
Computation and Language
The BEIR dataset is a large, heterogeneous benchmark for Information Retrieval (IR) in zero-shot settings, garnering considerable attention within the research community. However, BEIR and analogous datasets are predominantly restricted to the English language. Our objective is to establish extensive large-scale resources for IR in the Polish language, thereby advancing the research in this NLP area. In this work, inspired by mMARCO and Mr.~TyDi datasets, we translated all accessible open IR datasets into Polish, and we introduced the BEIR-PL benchmark -- a new benchmark which comprises 13 datasets, facilitating further development, training and evaluation of modern Polish language models for IR tasks. We executed an evaluation and comparison of numerous IR models on the newly introduced BEIR-PL benchmark. Furthermore, we publish pre-trained open IR models for Polish language,d marking a pioneering development in this field. Additionally, the evaluation revealed that BM25 achieved significantly lower scores for Polish than for English, which can be attributed to high inflection and intricate morphological structure of the Polish language. Finally, we trained various re-ranking models to enhance the BM25 retrieval, and we compared their performance to identify their unique characteristic features. To ensure accurate model comparisons, it is necessary to scrutinise individual results rather than to average across the entire benchmark. Thus, we thoroughly analysed the outcomes of IR models in relation to each individual data subset encompassed by the BEIR benchmark. The benchmark data is available at URL {\bf https://huggingface.co/clarin-knext}.
title BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2305.19840