LLMzSzŁ: a comprehensive LLM benchmark for Polish
Fuente:
arXiv
Saved in:
| Main Authors: | Jassem, Krzysztof, Ciesiółka, Michał, Graliński, Filip, Jabłoński, Piotr, Pokrywka, Jakub, Kubis, Marek, Jabłońska, Monika, Staruch, Ryszard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Polish-English medical knowledge transfer: A new benchmark and results
by: Grzybowski, Łukasz, et al.
Published: (2024)
by: Grzybowski, Łukasz, et al.
Published: (2024)
Two Approaches to Diachronic Normalization of Polish Texts
by: Dudzic, Kacper, et al.
Published: (2024)
by: Dudzic, Kacper, et al.
Published: (2024)
Oddballness: universal anomaly detection with language models
by: Graliński, Filip, et al.
Published: (2024)
by: Graliński, Filip, et al.
Published: (2024)
POLygraph: Polish Fake News Dataset
by: Dzienisiewicz, Daniel, et al.
Published: (2024)
by: Dzienisiewicz, Daniel, et al.
Published: (2024)
Adapting LLMs for Minimal-edit Grammatical Error Correction
by: Staruch, Ryszard, et al.
Published: (2025)
by: Staruch, Ryszard, et al.
Published: (2025)
Temporal Image Caption Retrieval Competition -- Description and Results
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Punctuation Prediction for Polish Texts using Transformers
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
Passage Retrieval of Polish Texts Using OKAPI BM25 and an Ensemble of Cross Encoders
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
GPT-4 passes most of the 297 written Polish Board Certification Examinations
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Tackling prediction tasks in relational databases with LLMs
by: Wydmuch, Marek, et al.
Published: (2024)
by: Wydmuch, Marek, et al.
Published: (2024)
ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation
by: Ciesiółka, Michał, et al.
Published: (2026)
by: Ciesiółka, Michał, et al.
Published: (2026)
Optimizing Retrieval-Augmented Generation of Medical Content for Spaced Repetition Learning
by: Kaczmarek, Jeremi I., et al.
Published: (2025)
by: Kaczmarek, Jeremi I., et al.
Published: (2025)
Evaluating Transformer Models for Suicide Risk Detection on Social Media
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Dynamic Boundary Time Warping for Sub-sequence Matching with Few Examples
by: Borchmann, Łukasz, et al.
Published: (2020)
by: Borchmann, Łukasz, et al.
Published: (2020)
CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs
by: Guttmann, Kamil, et al.
Published: (2026)
by: Guttmann, Kamil, et al.
Published: (2026)
CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
by: Zhang, Feng, et al.
Published: (2025)
by: Zhang, Feng, et al.
Published: (2025)
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)
by: Rybak, Piotr, et al.
Published: (2022)
A comprehensive study of LLM-based argument classification: from Llama through DeepSeek to GPT-5.2
by: Pietroń, Marcin, et al.
Published: (2026)
by: Pietroń, Marcin, et al.
Published: (2026)
A comprehensive study of LLM-based argument classification: from LLAMA through GPT-4o to Deepseek-R1
by: Pietroń, Marcin, et al.
Published: (2025)
by: Pietroń, Marcin, et al.
Published: (2025)
Long-Context Encoder Models for Polish Language Understanding
by: Dadas, Sławomir, et al.
Published: (2026)
by: Dadas, Sławomir, et al.
Published: (2026)
Inference Scaling for Bridging Retrieval and Augmented Generation
by: Lee, Youngwon, et al.
Published: (2024)
by: Lee, Youngwon, et al.
Published: (2024)
CORD: Balancing COnsistency and Rank Distillation for Robust Retrieval-Augmented Generation
by: Lee, Youngwon, et al.
Published: (2024)
by: Lee, Youngwon, et al.
Published: (2024)
ClonEval: An Open Voice Cloning Benchmark
by: Christop, Iwona, et al.
Published: (2025)
by: Christop, Iwona, et al.
Published: (2025)
Do Not Change Me: On Transferring Entities Without Modification in Neural Machine Translation -- a Multilingual Perspective
by: Wisniewski, Dawid, et al.
Published: (2025)
by: Wisniewski, Dawid, et al.
Published: (2025)
Seamlessly Integrating Tree-Based Positional Embeddings into Transformer Models for Source Code Representation
by: Bartkowiak, Patryk, et al.
Published: (2025)
by: Bartkowiak, Patryk, et al.
Published: (2025)
Large Language Models in Legislative Content Analysis: A Dataset from the Polish Parliament
by: Bryłkowski, Arkadiusz, et al.
Published: (2025)
by: Bryłkowski, Arkadiusz, et al.
Published: (2025)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
Cross-Family Speculative Decoding for Polish Language Models on Apple~Silicon: An Empirical Evaluation of Bielik~11B with UAG-Extended MLX-LM
by: Fonal, Krzysztof
Published: (2026)
by: Fonal, Krzysztof
Published: (2026)
Silver Retriever: Advancing Neural Passage Retrieval for Polish Question Answering
by: Rybak, Piotr, et al.
Published: (2023)
by: Rybak, Piotr, et al.
Published: (2023)
PL-MTEB: Polish Massive Text Embedding Benchmark
by: Poświata, Rafał, et al.
Published: (2024)
by: Poświata, Rafał, et al.
Published: (2024)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
by: Kubis, Marek, et al.
Published: (2025)
by: Kubis, Marek, et al.
Published: (2025)
ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT
by: Pokrywka, Mikołaj, et al.
Published: (2025)
by: Pokrywka, Mikołaj, et al.
Published: (2025)
Bielik-Q2-Sharp: A Comparative Study of Extreme 2-bit Quantization Methods for a Polish 11B Language Model
by: Prejzner, Jakub
Published: (2026)
by: Prejzner, Jakub
Published: (2026)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
by: Liu, Shuyu, et al.
Published: (2025)
by: Liu, Shuyu, et al.
Published: (2025)
The Polish Vocabulary Size Test: A Novel Adaptive Test for Receptive Vocabulary Assessment
by: Fokin, Danil, et al.
Published: (2025)
by: Fokin, Danil, et al.
Published: (2025)
Evaluating Polish linguistic and cultural competency in large language models
by: Dadas, Sławomir, et al.
Published: (2025)
by: Dadas, Sławomir, et al.
Published: (2025)
LLM-as-a-Judge is Bad, Based on AI Attempting the Exam Qualifying for the Member of the Polish National Board of Appeal
by: Karp, Michał, et al.
Published: (2025)
by: Karp, Michał, et al.
Published: (2025)
PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
by: Dadas, Sławomir, et al.
Published: (2024)
by: Dadas, Sławomir, et al.
Published: (2024)
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
by: Ociepa, Krzysztof, et al.
Published: (2024)
by: Ociepa, Krzysztof, et al.
Published: (2024)
PLLuM: A Family of Polish Large Language Models
by: Kocoń, Jan, et al.
Published: (2025)
by: Kocoń, Jan, et al.
Published: (2025)
Similar Items
-
Polish-English medical knowledge transfer: A new benchmark and results
by: Grzybowski, Łukasz, et al.
Published: (2024) -
Two Approaches to Diachronic Normalization of Polish Texts
by: Dudzic, Kacper, et al.
Published: (2024) -
Oddballness: universal anomaly detection with language models
by: Graliński, Filip, et al.
Published: (2024) -
POLygraph: Polish Fake News Dataset
by: Dzienisiewicz, Daniel, et al.
Published: (2024) -
Adapting LLMs for Minimal-edit Grammatical Error Correction
by: Staruch, Ryszard, et al.
Published: (2025)