AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zaghouani, Wajdi, Aldous, Kholoud K., Fejzullaj, Isra
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913169454137344
author Zaghouani, Wajdi
Aldous, Kholoud K.
Fejzullaj, Isra
author_facet Zaghouani, Wajdi
Aldous, Kholoud K.
Fejzullaj, Isra
contents Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety evaluation dataset for LLMs in Albanian, a linguistically distinct low-resource language with approximately 7.5 million speakers across Albania, Kosovo, North Macedonia, and the diaspora. The dataset contains 2,951 prompts spanning 11 safety categories, including self-harm, violence, racist content, child exploitation, and radicalization, with an average of 268 prompts per category. Each prompt is provided in Albanian with an English reference translation and a detailed category label. This resource addresses a significant gap in safety evaluation infrastruc-ture for low-resource languages and provides an essential benchmark for developing safer, more inclusive LLMs. The dataset will be provided upon request to support safety evaluation, fine-tuning, red-teaming, and guardrail development for Albanian-speaking communities.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26954
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
Zaghouani, Wajdi
Aldous, Kholoud K.
Fejzullaj, Isra
Computation and Language
Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety evaluation dataset for LLMs in Albanian, a linguistically distinct low-resource language with approximately 7.5 million speakers across Albania, Kosovo, North Macedonia, and the diaspora. The dataset contains 2,951 prompts spanning 11 safety categories, including self-harm, violence, racist content, child exploitation, and radicalization, with an average of 268 prompts per category. Each prompt is provided in Albanian with an English reference translation and a detailed category label. This resource addresses a significant gap in safety evaluation infrastruc-ture for low-resource languages and provides an essential benchmark for developing safer, more inclusive LLMs. The dataset will be provided upon request to support safety evaluation, fine-tuning, red-teaming, and guardrail development for Albanian-speaking communities.
title AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
topic Computation and Language
url https://arxiv.org/abs/2605.26954