CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Chaeyun, Lim, YongTaek, Kim, Kihyun, Kim, Junghwan, Kim, Minwoo
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918351307014144
author Kim, Chaeyun
Lim, YongTaek
Kim, Kihyun
Kim, Junghwan
Kim, Minwoo
author_facet Kim, Chaeyun
Lim, YongTaek
Kim, Kihyun
Kim, Junghwan
Kim, Minwoo
contents Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a critical blind spot in LLM safety evaluation. To address this gap, we introduce CAGE (Culturally Adaptive Generation), a framework that systematically adapts the adversarial intent of proven red-teaming prompts to new cultural contexts. At the core of CAGE is the Semantic Mold, a novel approach that disentangles a prompt's adversarial structure from its cultural content. This approach enables the modeling of realistic, localized threats rather than testing for simple jailbreaks. As a representative example, we demonstrate our framework by creating KoRSET, a Korean benchmark, which proves more effective at revealing vulnerabilities than direct translation baselines. CAGE offers a scalable solution for developing meaningful, context-aware safety benchmarks across diverse cultures. Our dataset and evaluation rubrics are publicly available at https://github.com/selectstar-ai/CAGE-paper. (WARNING: This paper contains model outputs that can be offensive in nature.)
format Preprint
id arxiv_https___arxiv_org_abs_2602_20170
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
Kim, Chaeyun
Lim, YongTaek
Kim, Kihyun
Kim, Junghwan
Kim, Minwoo
Computers and Society
Artificial Intelligence
Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a critical blind spot in LLM safety evaluation. To address this gap, we introduce CAGE (Culturally Adaptive Generation), a framework that systematically adapts the adversarial intent of proven red-teaming prompts to new cultural contexts. At the core of CAGE is the Semantic Mold, a novel approach that disentangles a prompt's adversarial structure from its cultural content. This approach enables the modeling of realistic, localized threats rather than testing for simple jailbreaks. As a representative example, we demonstrate our framework by creating KoRSET, a Korean benchmark, which proves more effective at revealing vulnerabilities than direct translation baselines. CAGE offers a scalable solution for developing meaningful, context-aware safety benchmarks across diverse cultures. Our dataset and evaluation rubrics are publicly available at https://github.com/selectstar-ai/CAGE-paper. (WARNING: This paper contains model outputs that can be offensive in nature.)
title CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2602.20170