PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dadas, Sławomir, Perełkiewicz, Michał, Poświata, Rafał
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909133320486912
author Dadas, Sławomir
Perełkiewicz, Michał
Poświata, Rafał
author_facet Dadas, Sławomir
Perełkiewicz, Michał
Poświata, Rafał
contents We present Polish Information Retrieval Benchmark (PIRB), a comprehensive evaluation framework encompassing 41 text information retrieval tasks for Polish. The benchmark incorporates existing datasets as well as 10 new, previously unpublished datasets covering diverse topics such as medicine, law, business, physics, and linguistics. We conduct an extensive evaluation of over 20 dense and sparse retrieval models, including the baseline models trained by us as well as other available Polish and multilingual methods. Finally, we introduce a three-step process for training highly effective language-specific retrievers, consisting of knowledge distillation, supervised fine-tuning, and building sparse-dense hybrid retrievers using a lightweight rescoring model. In order to validate our approach, we train new text encoders for Polish and compare their results with previously evaluated methods. Our dense models outperform the best solutions available to date, and the use of hybrid methods further improves their performance.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13350
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
Dadas, Sławomir
Perełkiewicz, Michał
Poświata, Rafał
Computation and Language
We present Polish Information Retrieval Benchmark (PIRB), a comprehensive evaluation framework encompassing 41 text information retrieval tasks for Polish. The benchmark incorporates existing datasets as well as 10 new, previously unpublished datasets covering diverse topics such as medicine, law, business, physics, and linguistics. We conduct an extensive evaluation of over 20 dense and sparse retrieval models, including the baseline models trained by us as well as other available Polish and multilingual methods. Finally, we introduce a three-step process for training highly effective language-specific retrievers, consisting of knowledge distillation, supervised fine-tuning, and building sparse-dense hybrid retrievers using a lightweight rescoring model. In order to validate our approach, we train new text encoders for Polish and compare their results with previously evaluated methods. Our dense models outperform the best solutions available to date, and the use of hybrid methods further improves their performance.
title PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
topic Computation and Language
url https://arxiv.org/abs/2402.13350