Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
Fuente:
arXiv
Saved in:
| Main Authors: | Wróbel, Krzysztof, Kowalski, Jan Maria, Surma, Jerzy, Ciuciura, Igor, Szymański, Maciej |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
by: Ociepa, Krzysztof, et al.
Published: (2024)
by: Ociepa, Krzysztof, et al.
Published: (2024)
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
by: Ociepa, Krzysztof, et al.
Published: (2026)
by: Ociepa, Krzysztof, et al.
Published: (2026)
Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language
by: Kinas, Remigiusz, et al.
Published: (2026)
by: Kinas, Remigiusz, et al.
Published: (2026)
Bielik 11B v3: Multilingual Large Language Model for European Languages
by: Ociepa, Krzysztof, et al.
Published: (2025)
by: Ociepa, Krzysztof, et al.
Published: (2025)
Bielik 11B v2 Technical Report
by: Ociepa, Krzysztof, et al.
Published: (2025)
by: Ociepa, Krzysztof, et al.
Published: (2025)
PL-Guard: Benchmarking Language Model Safety for Polish
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
Bielik v3 Small: Technical Report
by: Ociepa, Krzysztof, et al.
Published: (2025)
by: Ociepa, Krzysztof, et al.
Published: (2025)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
Experimentation in Content Moderation using RWKV
by: Yildirim, Umut, et al.
Published: (2024)
by: Yildirim, Umut, et al.
Published: (2024)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025)
by: Mehta, Manisha, et al.
Published: (2025)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
by: Zhang, Peiyan, et al.
Published: (2025)
by: Zhang, Peiyan, et al.
Published: (2025)
VertAttack: Taking advantage of Text Classifiers' horizontal vision
by: Rusert, Jonathan
Published: (2024)
by: Rusert, Jonathan
Published: (2024)
WHoW: A Cross-domain Approach for Analysing Conversation Moderation
by: Chen, Ming-Bin, et al.
Published: (2024)
by: Chen, Ming-Bin, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Sigmoid Head for Quality Estimation under Language Ambiguity
by: Dinh, Tu Anh, et al.
Published: (2026)
by: Dinh, Tu Anh, et al.
Published: (2026)
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
by: Boratyn, Daria, et al.
Published: (2026)
by: Boratyn, Daria, et al.
Published: (2026)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
by: Fukui, Hiroki
Published: (2026)
by: Fukui, Hiroki
Published: (2026)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025)
by: Xin, Yuan, et al.
Published: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2024)
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2024)
Identifying Fairness Issues in Automatically Generated Testing Content
by: Stowe, Kevin, et al.
Published: (2024)
by: Stowe, Kevin, et al.
Published: (2024)
KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
by: Song, Chenyang, et al.
Published: (2023)
by: Song, Chenyang, et al.
Published: (2023)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
Towards Human Understanding of Paraphrase Types in Large Language Models
by: Meier, Dominik, et al.
Published: (2024)
by: Meier, Dominik, et al.
Published: (2024)
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
by: Deng, Cheng, et al.
Published: (2025)
by: Deng, Cheng, et al.
Published: (2025)
Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models
by: Ghinassi, Iacopo, et al.
Published: (2024)
by: Ghinassi, Iacopo, et al.
Published: (2024)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
by: Lasbordes, Maxence, et al.
Published: (2025)
by: Lasbordes, Maxence, et al.
Published: (2025)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
by: Klöser, Lars, et al.
Published: (2024)
by: Klöser, Lars, et al.
Published: (2024)
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
by: Wen, Zhihao, et al.
Published: (2025)
by: Wen, Zhihao, et al.
Published: (2025)
Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation
by: Antypas, Dimosthenis, et al.
Published: (2024)
by: Antypas, Dimosthenis, et al.
Published: (2024)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
by: Kim, Songsoo, et al.
Published: (2025)
by: Kim, Songsoo, et al.
Published: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
by: Karinshak, Elise, et al.
Published: (2024)
by: Karinshak, Elise, et al.
Published: (2024)
Similar Items
-
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
by: Ociepa, Krzysztof, et al.
Published: (2024) -
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
by: Ociepa, Krzysztof, et al.
Published: (2026) -
Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language
by: Kinas, Remigiusz, et al.
Published: (2026) -
Bielik 11B v3: Multilingual Large Language Model for European Languages
by: Ociepa, Krzysztof, et al.
Published: (2025) -
Bielik 11B v2 Technical Report
by: Ociepa, Krzysztof, et al.
Published: (2025)