PL-MTEB: Polish Massive Text Embedding Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Poświata, Rafał, Dadas, Sławomir, Perełkiewicz, Michał |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
by: Dadas, Sławomir, et al.
Published: (2024)
by: Dadas, Sławomir, et al.
Published: (2024)
Evaluating Polish linguistic and cultural competency in large language models
by: Dadas, Sławomir, et al.
Published: (2025)
by: Dadas, Sławomir, et al.
Published: (2025)
SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
by: Perełkiewicz, Michał, et al.
Published: (2025)
by: Perełkiewicz, Michał, et al.
Published: (2025)
Long-Context Encoder Models for Polish Language Understanding
by: Dadas, Sławomir, et al.
Published: (2026)
by: Dadas, Sławomir, et al.
Published: (2026)
A Review of the Challenges with Massive Web-mined Corpora Used in Large Language Models Pre-Training
by: Perełkiewicz, Michał, et al.
Published: (2024)
by: Perełkiewicz, Michał, et al.
Published: (2024)
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
by: Poświata, Rafał, et al.
Published: (2025)
by: Poświata, Rafał, et al.
Published: (2025)
Assessing generalization capability of text ranking models in Polish
by: Dadas, Sławomir, et al.
Published: (2024)
by: Dadas, Sławomir, et al.
Published: (2024)
FinMTEB: Finance Massive Text Embedding Benchmark
by: Tang, Yixuan, et al.
Published: (2025)
by: Tang, Yixuan, et al.
Published: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025)
by: Pham, Loc, et al.
Published: (2025)
FaMTEB: Massive Text Embedding Benchmark in Persian Language
by: Zinvandi, Erfan, et al.
Published: (2025)
by: Zinvandi, Erfan, et al.
Published: (2025)
AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
by: Uemura, Kosei, et al.
Published: (2025)
by: Uemura, Kosei, et al.
Published: (2025)
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
by: Banar, Nikolay, et al.
Published: (2025)
by: Banar, Nikolay, et al.
Published: (2025)
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
by: Chung, Isaac, et al.
Published: (2025)
by: Chung, Isaac, et al.
Published: (2025)
PL-Guard: Benchmarking Language Model Safety for Polish
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
MMTEB: Massive Multilingual Text Embedding Benchmark
by: Enevoldsen, Kenneth, et al.
Published: (2025)
by: Enevoldsen, Kenneth, et al.
Published: (2025)
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
by: Ciancone, Mathieu, et al.
Published: (2024)
by: Ciancone, Mathieu, et al.
Published: (2024)
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
by: Wojtasik, Konrad, et al.
Published: (2023)
by: Wojtasik, Konrad, et al.
Published: (2023)
Massive Sound Embedding Benchmark (MSEB)
by: Heigold, Georg, et al.
Published: (2026)
by: Heigold, Georg, et al.
Published: (2026)
MIEB: Massive Image Embedding Benchmark
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
by: Cao, Hongliu
Published: (2024)
by: Cao, Hongliu
Published: (2024)
The Massive Legal Embedding Benchmark (MLEB)
by: Butler, Umar, et al.
Published: (2025)
by: Butler, Umar, et al.
Published: (2025)
MAEB: Massive Audio Embedding Benchmark
by: Assadi, Adnan El, et al.
Published: (2026)
by: Assadi, Adnan El, et al.
Published: (2026)
Two Approaches to Diachronic Normalization of Polish Texts
by: Dudzic, Kacper, et al.
Published: (2024)
by: Dudzic, Kacper, et al.
Published: (2024)
Punctuation Prediction for Polish Texts using Transformers
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service
by: Kołos, Anna, et al.
Published: (2023)
by: Kołos, Anna, et al.
Published: (2023)
Applying Text Embedding Models for Efficient Analysis in Labeled Property Graphs
by: Podstawski, Michal
Published: (2025)
by: Podstawski, Michal
Published: (2025)
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
by: Snegirev, Artem, et al.
Published: (2024)
by: Snegirev, Artem, et al.
Published: (2024)
PLLuM: A Family of Polish Large Language Models
by: Kocoń, Jan, et al.
Published: (2025)
by: Kocoń, Jan, et al.
Published: (2025)
German Text Embedding Clustering Benchmark
by: Wehrli, Silvan, et al.
Published: (2024)
by: Wehrli, Silvan, et al.
Published: (2024)
AI Text Detectors and the Misclassification of Slightly Polished Arabic Text
by: Almohaimeed, Saleh, et al.
Published: (2025)
by: Almohaimeed, Saleh, et al.
Published: (2025)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024)
by: Enevoldsen, Kenneth, et al.
Published: (2024)
IMGTB: A Framework for Machine-Generated Text Detection Benchmarking
by: Spiegel, Michal, et al.
Published: (2023)
by: Spiegel, Michal, et al.
Published: (2023)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
by: Yang, Lingyi, et al.
Published: (2023)
by: Yang, Lingyi, et al.
Published: (2023)
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
by: Omnilingual SONAR Team, et al.
Published: (2026)
by: Omnilingual SONAR Team, et al.
Published: (2026)
JFinTEB: Japanese Financial Text Embedding Benchmark
by: Suzuki, Masahiro, et al.
Published: (2026)
by: Suzuki, Masahiro, et al.
Published: (2026)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
by: Baek, Shaun, et al.
Published: (2025)
by: Baek, Shaun, et al.
Published: (2025)
ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain
by: Kasmaee, Ali Shiraee, et al.
Published: (2024)
by: Kasmaee, Ali Shiraee, et al.
Published: (2024)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
by: Luo, Yuxuan, et al.
Published: (2025)
by: Luo, Yuxuan, et al.
Published: (2025)
Similar Items
-
PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
by: Dadas, Sławomir, et al.
Published: (2024) -
Evaluating Polish linguistic and cultural competency in large language models
by: Dadas, Sławomir, et al.
Published: (2025) -
SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
by: Perełkiewicz, Michał, et al.
Published: (2025) -
Long-Context Encoder Models for Polish Language Understanding
by: Dadas, Sławomir, et al.
Published: (2026) -
A Review of the Challenges with Massive Web-mined Corpora Used in Large Language Models Pre-Training
by: Perełkiewicz, Michał, et al.
Published: (2024)