PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Dadas, Sławomir, Perełkiewicz, Michał, Poświata, Rafał |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PL-MTEB: Polish Massive Text Embedding Benchmark
by: Poświata, Rafał, et al.
Published: (2024)
by: Poświata, Rafał, et al.
Published: (2024)
Evaluating Polish linguistic and cultural competency in large language models
by: Dadas, Sławomir, et al.
Published: (2025)
by: Dadas, Sławomir, et al.
Published: (2025)
SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
by: Perełkiewicz, Michał, et al.
Published: (2025)
by: Perełkiewicz, Michał, et al.
Published: (2025)
Long-Context Encoder Models for Polish Language Understanding
by: Dadas, Sławomir, et al.
Published: (2026)
by: Dadas, Sławomir, et al.
Published: (2026)
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
by: Poświata, Rafał, et al.
Published: (2025)
by: Poświata, Rafał, et al.
Published: (2025)
A Review of the Challenges with Massive Web-mined Corpora Used in Large Language Models Pre-Training
by: Perełkiewicz, Michał, et al.
Published: (2024)
by: Perełkiewicz, Michał, et al.
Published: (2024)
Assessing generalization capability of text ranking models in Polish
by: Dadas, Sławomir, et al.
Published: (2024)
by: Dadas, Sławomir, et al.
Published: (2024)
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
by: Wojtasik, Konrad, et al.
Published: (2023)
by: Wojtasik, Konrad, et al.
Published: (2023)
Passage Retrieval of Polish Texts Using OKAPI BM25 and an Ensemble of Cross Encoders
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
PLLuM: A Family of Polish Large Language Models
by: Kocoń, Jan, et al.
Published: (2025)
by: Kocoń, Jan, et al.
Published: (2025)
Two Approaches to Diachronic Normalization of Polish Texts
by: Dudzic, Kacper, et al.
Published: (2024)
by: Dudzic, Kacper, et al.
Published: (2024)
Punctuation Prediction for Polish Texts using Transformers
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment
by: Luo, Kun, et al.
Published: (2024)
by: Luo, Kun, et al.
Published: (2024)
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
by: Song, Jonghyun, et al.
Published: (2025)
by: Song, Jonghyun, et al.
Published: (2025)
IMGTB: A Framework for Machine-Generated Text Detection Benchmarking
by: Spiegel, Michal, et al.
Published: (2023)
by: Spiegel, Michal, et al.
Published: (2023)
Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems
by: Zhuang, Shengyao, et al.
Published: (2024)
by: Zhuang, Shengyao, et al.
Published: (2024)
Dense Passage Retrieval: Is it Retrieving?
by: Reichman, Benjamin, et al.
Published: (2024)
by: Reichman, Benjamin, et al.
Published: (2024)
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
by: Mahalingam, Aakash, et al.
Published: (2024)
by: Mahalingam, Aakash, et al.
Published: (2024)
Silver Retriever: Advancing Neural Passage Retrieval for Polish Question Answering
by: Rybak, Piotr, et al.
Published: (2023)
by: Rybak, Piotr, et al.
Published: (2023)
AI Text Detectors and the Misclassification of Slightly Polished Arabic Text
by: Almohaimeed, Saleh, et al.
Published: (2025)
by: Almohaimeed, Saleh, et al.
Published: (2025)
QAEA-DR: A Unified Text Augmentation Framework for Dense Retrieval
by: Tan, Hongming, et al.
Published: (2024)
by: Tan, Hongming, et al.
Published: (2024)
A Reasoning-Focused Legal Retrieval Benchmark
by: Zheng, Lucia, et al.
Published: (2025)
by: Zheng, Lucia, et al.
Published: (2025)
Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense Retrieval
by: Liu, Yuxiang, et al.
Published: (2025)
by: Liu, Yuxiang, et al.
Published: (2025)
A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts
by: Rayo, Jhon, et al.
Published: (2025)
by: Rayo, Jhon, et al.
Published: (2025)
A Method for Handling Negative Similarities in Explainable Graph Spectral Clustering of Text Documents -- Extended Version
by: Kłopotek, Mieczysław A., et al.
Published: (2025)
by: Kłopotek, Mieczysław A., et al.
Published: (2025)
Boolean-aware Attention for Dense Retrieval
by: Mai, Quan, et al.
Published: (2025)
by: Mai, Quan, et al.
Published: (2025)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
by: Yang, Lingyi, et al.
Published: (2023)
by: Yang, Lingyi, et al.
Published: (2023)
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
by: Li, Xiangyang, et al.
Published: (2024)
by: Li, Xiangyang, et al.
Published: (2024)
Cohort Retrieval using Dense Passage Retrieval
by: Jadhav, Pranav
Published: (2025)
by: Jadhav, Pranav
Published: (2025)
SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Scaling Laws For Dense Retrieval
by: Fang, Yan, et al.
Published: (2024)
by: Fang, Yan, et al.
Published: (2024)
Multilingual Entity Linking Using Dense Retrieval
by: Farhan, Dominik
Published: (2024)
by: Farhan, Dominik
Published: (2024)
Dynamic Injection of Entity Knowledge into Dense Retrievers
by: Yamada, Ikuya, et al.
Published: (2025)
by: Yamada, Ikuya, et al.
Published: (2025)
DReSD: Dense Retrieval for Speculative Decoding
by: Gritta, Milan, et al.
Published: (2025)
by: Gritta, Milan, et al.
Published: (2025)
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
by: Lyu, Yuanjie, et al.
Published: (2024)
by: Lyu, Yuanjie, et al.
Published: (2024)
Polish-English medical knowledge transfer: A new benchmark and results
by: Grzybowski, Łukasz, et al.
Published: (2024)
by: Grzybowski, Łukasz, et al.
Published: (2024)
HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs
by: Li, Shujie, et al.
Published: (2025)
by: Li, Shujie, et al.
Published: (2025)
PL-Guard: Benchmarking Language Model Safety for Polish
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
by: Pereira, Jayr, et al.
Published: (2026)
by: Pereira, Jayr, et al.
Published: (2026)
Similar Items
-
PL-MTEB: Polish Massive Text Embedding Benchmark
by: Poświata, Rafał, et al.
Published: (2024) -
Evaluating Polish linguistic and cultural competency in large language models
by: Dadas, Sławomir, et al.
Published: (2025) -
SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
by: Perełkiewicz, Michał, et al.
Published: (2025) -
Long-Context Encoder Models for Polish Language Understanding
by: Dadas, Sławomir, et al.
Published: (2026) -
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
by: Poświata, Rafał, et al.
Published: (2025)