KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zaghouani, Wajdi, Ibrahim, Shimaa Amer, Muratbek, Aruzhan, Zhakenov, Olzhasbek, Akhmetzhanova, Adiya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
Cultural Adaptation in Large Language Models for Political Discourse
by: Zaghouani, Wajdi
Published: (2026)
by: Zaghouani, Wajdi
Published: (2026)
Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems
by: Zaghouani, Wajdi
Published: (2026)
by: Zaghouani, Wajdi
Published: (2026)
Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities
by: Zaghouani, Wajdi
Published: (2026)
by: Zaghouani, Wajdi
Published: (2026)
Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic
by: Zaghouani, Wajdi, et al.
Published: (2025)
by: Zaghouani, Wajdi, et al.
Published: (2025)
AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse
by: Sharqawi, Esra'a, et al.
Published: (2026)
by: Sharqawi, Esra'a, et al.
Published: (2026)
Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse
by: Al-Athba, Aisha Ali, et al.
Published: (2026)
by: Al-Athba, Aisha Ali, et al.
Published: (2026)
An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
by: Zaghouani, Wajdi, et al.
Published: (2025)
by: Zaghouani, Wajdi, et al.
Published: (2025)
Chinese Offensive Language Detection:Current Status and Future Directions
by: Xiao, Yunze, et al.
Published: (2024)
by: Xiao, Yunze, et al.
Published: (2024)
SozKZ: Training Efficient Small Language Models for Kazakh from Scratch
by: Tukenov, Saken
Published: (2026)
by: Tukenov, Saken
Published: (2026)
ArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
MARSAD: A Multi-Functional Tool for Real-Time Social Media Analysis
by: Biswas, Md. Rafiul, et al.
Published: (2025)
by: Biswas, Md. Rafiul, et al.
Published: (2025)
MemeMind at ArAIEval Shared Task: Spotting Persuasive Spans in Arabic Text with Persuasion Techniques Identification
by: Biswas, Md Rafiul, et al.
Published: (2024)
by: Biswas, Md Rafiul, et al.
Published: (2024)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
by: Nakanishi, Akito, et al.
Published: (2025)
by: Nakanishi, Akito, et al.
Published: (2025)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
by: Goloburda, Maiya, et al.
Published: (2025)
by: Goloburda, Maiya, et al.
Published: (2025)
Nullpointer at CheckThat! 2024: Identifying Subjectivity from Multilingual Text Sequence
by: Biswas, Md. Rafiul, et al.
Published: (2024)
by: Biswas, Md. Rafiul, et al.
Published: (2024)
Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models
by: Long, Do Xuan, et al.
Published: (2024)
by: Long, Do Xuan, et al.
Published: (2024)
SafetyBench: Evaluating the Safety of Large Language Models
by: Zhang, Zhexin, et al.
Published: (2023)
by: Zhang, Zhexin, et al.
Published: (2023)
From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Model
by: Hu, Mengya, et al.
Published: (2026)
by: Hu, Mengya, et al.
Published: (2026)
ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
by: Tang, Peiyuan, et al.
Published: (2025)
by: Tang, Peiyuan, et al.
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
LongSafety: Evaluating Long-Context Safety of Large Language Models
by: Lu, Yida, et al.
Published: (2025)
by: Lu, Yida, et al.
Published: (2025)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models
by: Lutz, Marlene, et al.
Published: (2025)
by: Lutz, Marlene, et al.
Published: (2025)
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)
by: Li, Xingxuan, et al.
Published: (2022)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Large Language Model Prompt Datasets: An In-depth Analysis and Insights
by: Zhang, Yuanming, et al.
Published: (2025)
by: Zhang, Yuanming, et al.
Published: (2025)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
by: Mou, Yutao, et al.
Published: (2024)
by: Mou, Yutao, et al.
Published: (2024)
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
by: Torres-Fonseca, Josue, et al.
Published: (2026)
by: Torres-Fonseca, Josue, et al.
Published: (2026)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
PromptExp: Multi-granularity Prompt Explanation of Large Language Models
by: Dong, Ximing, et al.
Published: (2024)
by: Dong, Ximing, et al.
Published: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
Gnosis Prompt: Adaptive Safety Layer for Large Language Models
by: González Medina, Claudio
Published: (2025)
by: González Medina, Claudio
Published: (2025)
Similar Items
-
AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
by: Zaghouani, Wajdi, et al.
Published: (2026) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024) -
ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination
by: Zaghouani, Wajdi, et al.
Published: (2026) -
JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media
by: Zaghouani, Wajdi, et al.
Published: (2026) -
Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus
by: Zaghouani, Wajdi, et al.
Published: (2026)