PL-Guard: Benchmarking Language Model Safety for Polish
Fuente:
arXiv
Saved in:
| Main Authors: | Krasnodębska, Aleksandra, Seweryn, Karolina, Łukasik, Szymon, Kusa, Wojciech |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
by: Boratyn, Daria, et al.
Published: (2026)
by: Boratyn, Daria, et al.
Published: (2026)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
by: Pachinger, Pia, et al.
Published: (2024)
by: Pachinger, Pia, et al.
Published: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
by: Zhang, Peiyan, et al.
Published: (2025)
by: Zhang, Peiyan, et al.
Published: (2025)
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
by: Gupta, Abhay, et al.
Published: (2025)
by: Gupta, Abhay, et al.
Published: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language
by: Kinas, Remigiusz, et al.
Published: (2026)
by: Kinas, Remigiusz, et al.
Published: (2026)
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
by: Ociepa, Krzysztof, et al.
Published: (2024)
by: Ociepa, Krzysztof, et al.
Published: (2024)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
by: Gaim, Fitsum, et al.
Published: (2025)
by: Gaim, Fitsum, et al.
Published: (2025)
EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation
by: Yang, Pei, et al.
Published: (2026)
by: Yang, Pei, et al.
Published: (2026)
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
by: Ociepa, Krzysztof, et al.
Published: (2026)
by: Ociepa, Krzysztof, et al.
Published: (2026)
LLM generated responses to mitigate the impact of hate speech
by: Podolak, Jakub, et al.
Published: (2023)
by: Podolak, Jakub, et al.
Published: (2023)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Precise Length Control in Large Language Models
by: Butcher, Bradley, et al.
Published: (2024)
by: Butcher, Bradley, et al.
Published: (2024)
What Drives Performance in Multilingual Language Models?
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
Large Language Models for Biomedical Article Classification
by: Proboszcz, Jakub, et al.
Published: (2026)
by: Proboszcz, Jakub, et al.
Published: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Strategy Adaptation in Large Language Model Werewolf Agents
by: Nakamori, Fuya, et al.
Published: (2025)
by: Nakamori, Fuya, et al.
Published: (2025)
Socially Responsible Data for Large Multilingual Language Models
by: Smart, Andrew, et al.
Published: (2024)
by: Smart, Andrew, et al.
Published: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
ADALog: Adaptive Unsupervised Anomaly detection in Logs with Self-attention Masked Language Model
by: Pospieszny, Przemek, et al.
Published: (2025)
by: Pospieszny, Przemek, et al.
Published: (2025)
Qomhra: A Bilingual Irish and English Large Language Model
by: McInerney, Joseph, et al.
Published: (2025)
by: McInerney, Joseph, et al.
Published: (2025)
Dialect Normalization using Large Language Models and Morphological Rules
by: Dimakis, Antonios, et al.
Published: (2025)
by: Dimakis, Antonios, et al.
Published: (2025)
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
by: Rezaeimanesh, Sara, et al.
Published: (2024)
by: Rezaeimanesh, Sara, et al.
Published: (2024)
Towards Human Understanding of Paraphrase Types in Large Language Models
by: Meier, Dominik, et al.
Published: (2024)
by: Meier, Dominik, et al.
Published: (2024)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
Task Contamination: Language Models May Not Be Few-Shot Anymore
by: Li, Changmao, et al.
Published: (2023)
by: Li, Changmao, et al.
Published: (2023)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
by: Peláez-González, Carlos, et al.
Published: (2025)
by: Peláez-González, Carlos, et al.
Published: (2025)
Linguistic Interpretability of Transformer-based Language Models: a systematic review
by: López-Otal, Miguel, et al.
Published: (2025)
by: López-Otal, Miguel, et al.
Published: (2025)
KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
by: Zhao, Yong, et al.
Published: (2025)
by: Zhao, Yong, et al.
Published: (2025)
Similar Items
-
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026) -
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
by: Boratyn, Daria, et al.
Published: (2026) -
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025) -
AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
by: Pachinger, Pia, et al.
Published: (2024) -
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
by: Pan, Leyi, et al.
Published: (2025)