Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yujia, Hee, Ming Shan, Nakov, Preslav, Lee, Roy Ka-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912600810323968
author Hu, Yujia
Hee, Ming Shan
Nakov, Preslav
Lee, Roy Ka-Wei
author_facet Hu, Yujia
Hee, Ming Shan
Nakov, Preslav
Lee, Roy Ka-Wei
contents The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a novel dataset and evaluation framework for benchmarking LLM safety in Singapore's diverse linguistic context, including Singlish, Chinese, Malay, and Tamil. SGToxicGuard adopts a red-teaming approach to systematically probe LLM vulnerabilities in three real-world scenarios: \textit{conversation}, \textit{question-answering}, and \textit{content composition}. We conduct extensive experiments with state-of-the-art multilingual LLMs, and the results uncover critical gaps in their safety guardrails. By offering actionable insights into cultural sensitivity and toxicity mitigation, we lay the foundation for safer and more inclusive AI systems in linguistically diverse environments.\footnote{Link to the dataset: https://github.com/Social-AI-Studio/SGToxicGuard.} \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}
format Preprint
id arxiv_https___arxiv_org_abs_2509_15260
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
Hu, Yujia
Hee, Ming Shan
Nakov, Preslav
Lee, Roy Ka-Wei
Computation and Language
The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a novel dataset and evaluation framework for benchmarking LLM safety in Singapore's diverse linguistic context, including Singlish, Chinese, Malay, and Tamil. SGToxicGuard adopts a red-teaming approach to systematically probe LLM vulnerabilities in three real-world scenarios: \textit{conversation}, \textit{question-answering}, and \textit{content composition}. We conduct extensive experiments with state-of-the-art multilingual LLMs, and the results uncover critical gaps in their safety guardrails. By offering actionable insights into cultural sensitivity and toxicity mitigation, we lay the foundation for safer and more inclusive AI systems in linguistically diverse environments.\footnote{Link to the dataset: https://github.com/Social-AI-Studio/SGToxicGuard.} \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}
title Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2509.15260