TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chu, Hua-Rong, Wang, Kuan-Chun, Huang, Yao-Te
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918453018886144
author Chu, Hua-Rong
Wang, Kuan-Chun
Huang, Yao-Te
author_facet Chu, Hua-Rong
Wang, Kuan-Chun
Huang, Yao-Te
contents Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). However, existing research lacks consideration of nuances across linguistic and cultural contexts, resulting in a gap between reported performance and in-the-wild effectiveness. To address this issue, this paper proposes an approach to optimize guardrail models for a designated linguistic context by leveraging a curated dataset tailored to local linguistic characteristics, targeting the Taiwan linguistic context as a representative example of localized deployment challenges. The proposed approach yields TWGuard, a linguistic context-optimized guardrail model that achieves a huge gain (+0.289 in F1) compared to the foundation model and significantly outperforms the strongest baseline in practical use (-0.037 in false positive rate, a 94.9\% reduction). Together, this work lays a foundation for regional communities to establish AI safety standards grounded in their own linguistic contexts, rather than accepting boundaries imposed by dominant languages. The inadequacy of the latter is reconfirmed by our findings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16542
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
Chu, Hua-Rong
Wang, Kuan-Chun
Huang, Yao-Te
Cryptography and Security
Computation and Language
Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). However, existing research lacks consideration of nuances across linguistic and cultural contexts, resulting in a gap between reported performance and in-the-wild effectiveness. To address this issue, this paper proposes an approach to optimize guardrail models for a designated linguistic context by leveraging a curated dataset tailored to local linguistic characteristics, targeting the Taiwan linguistic context as a representative example of localized deployment challenges. The proposed approach yields TWGuard, a linguistic context-optimized guardrail model that achieves a huge gain (+0.289 in F1) compared to the foundation model and significantly outperforms the strongest baseline in practical use (-0.037 in false positive rate, a 94.9\% reduction). Together, this work lays a foundation for regional communities to establish AI safety standards grounded in their own linguistic contexts, rather than accepting boundaries imposed by dominant languages. The inadequacy of the latter is reconfirmed by our findings.
title TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2604.16542