SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tasawong, Panuthep, Ngui, Jian Gang, Aji, Alham Fikri, Cohn, Trevor, Limkonchotiwat, Peerat
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909945984712704
author Tasawong, Panuthep
Ngui, Jian Gang
Aji, Alham Fikri
Cohn, Trevor
Limkonchotiwat, Peerat
author_facet Tasawong, Panuthep
Ngui, Jian Gang
Aji, Alham Fikri
Cohn, Trevor
Limkonchotiwat, Peerat
contents Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Existing multilingual safety benchmarks often rely on machine-translated English data, which fails to capture nuances in low-resource languages. Southeast Asian (SEA) languages are underrepresented despite the region's linguistic diversity and unique safety concerns, from culturally sensitive political speech to region-specific misinformation. Addressing these gaps requires benchmarks that are natively authored to reflect local norms and harm scenarios. We introduce SEA-SafeguardBench, the first human-verified safety benchmark for SEA, covering eight languages, 21,640 samples, across three subsets: general, in-the-wild, and content generation. The experimental results from our benchmark demonstrate that even state-of-the-art LLMs and guardrails are challenged by SEA cultural and harm scenarios and underperform when compared to English texts.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
Tasawong, Panuthep
Ngui, Jian Gang
Aji, Alham Fikri
Cohn, Trevor
Limkonchotiwat, Peerat
Computation and Language
Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Existing multilingual safety benchmarks often rely on machine-translated English data, which fails to capture nuances in low-resource languages. Southeast Asian (SEA) languages are underrepresented despite the region's linguistic diversity and unique safety concerns, from culturally sensitive political speech to region-specific misinformation. Addressing these gaps requires benchmarks that are natively authored to reflect local norms and harm scenarios. We introduce SEA-SafeguardBench, the first human-verified safety benchmark for SEA, covering eight languages, 21,640 samples, across three subsets: general, in-the-wild, and content generation. The experimental results from our benchmark demonstrate that even state-of-the-art LLMs and guardrails are challenged by SEA cultural and harm scenarios and underperform when compared to English texts.
title SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
topic Computation and Language
url https://arxiv.org/abs/2512.05501