PL-Guard: Benchmarking Language Model Safety for Polish

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Krasnodębska, Aleksandra, Seweryn, Karolina, Łukasik, Szymon, Kusa, Wojciech
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911014816055296
author Krasnodębska, Aleksandra
Seweryn, Karolina
Łukasik, Szymon
Kusa, Wojciech
author_facet Krasnodębska, Aleksandra
Seweryn, Karolina
Łukasik, Szymon
Kusa, Wojciech
contents Despite increasing efforts to ensure the safety of large language models (LLMs), most existing safety assessments and moderation tools remain heavily biased toward English and other high-resource languages, leaving majority of global languages underexamined. To address this gap, we introduce a manually annotated benchmark dataset for language model safety classification in Polish. We also create adversarially perturbed variants of these samples designed to challenge model robustness. We conduct a series of experiments to evaluate LLM-based and classifier-based models of varying sizes and architectures. Specifically, we fine-tune three models: Llama-Guard-3-8B, a HerBERT-based classifier (a Polish BERT derivative), and PLLuM, a Polish-adapted Llama-8B model. We train these models using different combinations of annotated data and evaluate their performance, comparing it against publicly available guard models. Results demonstrate that the HerBERT-based classifier achieves the highest overall performance, particularly under adversarial conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PL-Guard: Benchmarking Language Model Safety for Polish
Krasnodębska, Aleksandra
Seweryn, Karolina
Łukasik, Szymon
Kusa, Wojciech
Computation and Language
I.2.7
Despite increasing efforts to ensure the safety of large language models (LLMs), most existing safety assessments and moderation tools remain heavily biased toward English and other high-resource languages, leaving majority of global languages underexamined. To address this gap, we introduce a manually annotated benchmark dataset for language model safety classification in Polish. We also create adversarially perturbed variants of these samples designed to challenge model robustness. We conduct a series of experiments to evaluate LLM-based and classifier-based models of varying sizes and architectures. Specifically, we fine-tune three models: Llama-Guard-3-8B, a HerBERT-based classifier (a Polish BERT derivative), and PLLuM, a Polish-adapted Llama-8B model. We train these models using different combinations of annotated data and evaluate their performance, comparing it against publicly available guard models. Results demonstrate that the HerBERT-based classifier achieves the highest overall performance, particularly under adversarial conditions.
title PL-Guard: Benchmarking Language Model Safety for Polish
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2506.16322