skLEP: A Slovak General Language Understanding Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908424046903296 |
|---|---|
| author | Šuppa, Marek Ridzik, Andrej Hládek, Daniel Javůrek, Tomáš Ondrejová, Viktória Sásiková, Kristína Tamajka, Martin Šimko, Marián |
| author_facet | Šuppa, Marek Ridzik, Andrej Hládek, Daniel Javůrek, Tomáš Ondrejová, Viktória Sásiková, Kristína Tamajka, Martin Šimko, Marián |
| contents | In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP to encompass nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets tailored for Slovak and meticulously translated established English NLU resources. Within this paper, we also present the first systematic and extensive evaluation of a wide array of Slovak-specific, multilingual, and English pre-trained language models using the skLEP tasks. Finally, we also release the complete benchmark data, an open-source toolkit facilitating both fine-tuning and evaluation of models, and a public leaderboard at https://github.com/slovak-nlp/sklep in the hopes of fostering reproducibility and drive future research in Slovak NLU. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21508 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | skLEP: A Slovak General Language Understanding Benchmark Šuppa, Marek Ridzik, Andrej Hládek, Daniel Javůrek, Tomáš Ondrejová, Viktória Sásiková, Kristína Tamajka, Martin Šimko, Marián Computation and Language Artificial Intelligence Information Retrieval Machine Learning 68T50 I.2.7 In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP to encompass nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets tailored for Slovak and meticulously translated established English NLU resources. Within this paper, we also present the first systematic and extensive evaluation of a wide array of Slovak-specific, multilingual, and English pre-trained language models using the skLEP tasks. Finally, we also release the complete benchmark data, an open-source toolkit facilitating both fine-tuning and evaluation of models, and a public leaderboard at https://github.com/slovak-nlp/sklep in the hopes of fostering reproducibility and drive future research in Slovak NLU. |
| title | skLEP: A Slovak General Language Understanding Benchmark |
| topic | Computation and Language Artificial Intelligence Information Retrieval Machine Learning 68T50 I.2.7 |
| url | https://arxiv.org/abs/2506.21508 |