Saved in:
| Main Authors: | Grzybowski, Łukasz, Pokrywka, Jakub, Ciesiółka, Michał, Kaczmarek, Jeremi I., Kubis, Marek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.00559 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMzSzŁ: a comprehensive LLM benchmark for Polish
by: Jassem, Krzysztof, et al.
Published: (2025)
by: Jassem, Krzysztof, et al.
Published: (2025)
GPT-4 passes most of the 297 written Polish Board Certification Examinations
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Optimizing Retrieval-Augmented Generation of Medical Content for Spaced Repetition Learning
by: Kaczmarek, Jeremi I., et al.
Published: (2025)
by: Kaczmarek, Jeremi I., et al.
Published: (2025)
Passage Retrieval of Polish Texts Using OKAPI BM25 and an Ensemble of Cross Encoders
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
Evaluating Transformer Models for Suicide Risk Detection on Social Media
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Punctuation Prediction for Polish Texts using Transformers
by: Pokrywka, Jakub
Published: (2024)
by: Pokrywka, Jakub
Published: (2024)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
by: Kubis, Marek, et al.
Published: (2025)
by: Kubis, Marek, et al.
Published: (2025)
ADMEDTAGGER: an annotation framework for distillation of expert knowledge for the Polish medical language
by: Górski, Franciszek, et al.
Published: (2025)
by: Górski, Franciszek, et al.
Published: (2025)
Bielik-Q2-Sharp: A Comparative Study of Extreme 2-bit Quantization Methods for a Polish 11B Language Model
by: Prejzner, Jakub
Published: (2026)
by: Prejzner, Jakub
Published: (2026)
PLLuM: A Family of Polish Large Language Models
by: Kocoń, Jan, et al.
Published: (2025)
by: Kocoń, Jan, et al.
Published: (2025)
Chasing COMET: Leveraging Minimum Bayes Risk Decoding for Self-Improving Machine Translation
by: Guttmann, Kamil, et al.
Published: (2024)
by: Guttmann, Kamil, et al.
Published: (2024)
Self-training Language Models for Arithmetic Reasoning
by: Kadlčík, Marek, et al.
Published: (2024)
by: Kadlčík, Marek, et al.
Published: (2024)
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Polish-ASTE: Aspect-Sentiment Triplet Extraction Datasets for Polish
by: Lango, Marta, et al.
Published: (2025)
by: Lango, Marta, et al.
Published: (2025)
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
by: Ociepa, Krzysztof, et al.
Published: (2024)
by: Ociepa, Krzysztof, et al.
Published: (2024)
Concept-aware Data Construction Improves In-context Learning of Language Models
by: Štefánik, Michal, et al.
Published: (2024)
by: Štefánik, Michal, et al.
Published: (2024)
StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
by: Ochab, Jeremi K., et al.
Published: (2025)
by: Ochab, Jeremi K., et al.
Published: (2025)
Key ingredients for effective zero-shot cross-lingual knowledge transfer in generative tasks
by: Chirkova, Nadezhda, et al.
Published: (2024)
by: Chirkova, Nadezhda, et al.
Published: (2024)
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
by: Ociepa, Krzysztof, et al.
Published: (2026)
by: Ociepa, Krzysztof, et al.
Published: (2026)
Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models
by: Mikula, Lukáš, et al.
Published: (2023)
by: Mikula, Lukáš, et al.
Published: (2023)
Stylometry recognizes human and LLM-generated texts in short samples
by: Przystalski, Karol, et al.
Published: (2025)
by: Przystalski, Karol, et al.
Published: (2025)
How new data permeates LLM knowledge and how to dilute it
by: Sun, Chen, et al.
Published: (2025)
by: Sun, Chen, et al.
Published: (2025)
Cross-lingual Named Entity Corpus for Slavic Languages
by: Piskorski, Jakub, et al.
Published: (2024)
by: Piskorski, Jakub, et al.
Published: (2024)
Collaboratively adding new knowledge to an LLM
by: Lee, Rhui Dih, et al.
Published: (2024)
by: Lee, Rhui Dih, et al.
Published: (2024)
Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language
by: Kinas, Remigiusz, et al.
Published: (2026)
by: Kinas, Remigiusz, et al.
Published: (2026)
In Case You Missed It: ARC 'Challenge' Is Not That Challenging
by: Borchmann, Łukasz
Published: (2024)
by: Borchmann, Łukasz
Published: (2024)
Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
by: Junczyk, Michał
Published: (2024)
by: Junczyk, Michał
Published: (2024)
Suvach -- Generated Hindi QA benchmark
by: Narayanan, Vaishak, et al.
Published: (2024)
by: Narayanan, Vaishak, et al.
Published: (2024)
Two Approaches to Diachronic Normalization of Polish Texts
by: Dudzic, Kacper, et al.
Published: (2024)
by: Dudzic, Kacper, et al.
Published: (2024)
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025)
by: Stefaniak, Maciej, et al.
Published: (2025)
AI Text Detectors and the Misclassification of Slightly Polished Arabic Text
by: Almohaimeed, Saleh, et al.
Published: (2025)
by: Almohaimeed, Saleh, et al.
Published: (2025)
ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation
by: Ciesiółka, Michał, et al.
Published: (2026)
by: Ciesiółka, Michał, et al.
Published: (2026)
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
POLygraph: Polish Fake News Dataset
by: Dzienisiewicz, Daniel, et al.
Published: (2024)
by: Dzienisiewicz, Daniel, et al.
Published: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Larger models yield better results? Streamlined severity classification of ADHD-related concerns using BERT-based knowledge distillation
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
by: Gupta, Kshitij
Published: (2025)
by: Gupta, Kshitij
Published: (2025)
Similar Items
-
LLMzSzŁ: a comprehensive LLM benchmark for Polish
by: Jassem, Krzysztof, et al.
Published: (2025) -
GPT-4 passes most of the 297 written Polish Board Certification Examinations
by: Pokrywka, Jakub, et al.
Published: (2024) -
Optimizing Retrieval-Augmented Generation of Medical Content for Spaced Repetition Learning
by: Kaczmarek, Jeremi I., et al.
Published: (2025) -
Passage Retrieval of Polish Texts Using OKAPI BM25 and an Ensemble of Cross Encoders
by: Pokrywka, Jakub
Published: (2024) -
Evaluating Transformer Models for Suicide Risk Detection on Social Media
by: Pokrywka, Jakub, et al.
Published: (2024)