Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoo, Haneul, Yang, Yongjin, Lee, Hwaran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
von: Yang, Yongjin, et al.
Veröffentlicht: (2024)
von: Yang, Yongjin, et al.
Veröffentlicht: (2024)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
von: Yoo, Haneul, et al.
Veröffentlicht: (2025)
von: Yoo, Haneul, et al.
Veröffentlicht: (2025)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
von: Winata, Genta Indra, et al.
Veröffentlicht: (2026)
von: Winata, Genta Indra, et al.
Veröffentlicht: (2026)
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
von: Xie, Peng, et al.
Veröffentlicht: (2025)
von: Xie, Peng, et al.
Veröffentlicht: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
TroubleLLM: Align to Red Team Expert
von: Xu, Zhuoer, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoer, et al.
Veröffentlicht: (2024)
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
von: Asano, Shunta, et al.
Veröffentlicht: (2026)
von: Asano, Shunta, et al.
Veröffentlicht: (2026)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
von: Feier, Andrei Marian, et al.
Veröffentlicht: (2026)
von: Feier, Andrei Marian, et al.
Veröffentlicht: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
From National Curricula to Cultural Awareness: Constructing Open-Ended Culture-Specific Question Answering Dataset
von: Yoo, Haneul, et al.
Veröffentlicht: (2026)
von: Yoo, Haneul, et al.
Veröffentlicht: (2026)
OLA: Output Language Alignment in Code-Switched LLM Interactions
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations
von: Raheja, Tarun, et al.
Veröffentlicht: (2024)
von: Raheja, Tarun, et al.
Veröffentlicht: (2024)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
von: Bu, Yuyan, et al.
Veröffentlicht: (2026)
von: Bu, Yuyan, et al.
Veröffentlicht: (2026)
Parsing the Switch: LLM-Based UD Annotation for Complex Code-Switched and Low-Resource Languages
von: Kellert, Olga, et al.
Veröffentlicht: (2025)
von: Kellert, Olga, et al.
Veröffentlicht: (2025)
Exploring Straightforward Conversational Red-Teaming
von: Kour, George, et al.
Veröffentlicht: (2024)
von: Kour, George, et al.
Veröffentlicht: (2024)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
von: Yan, Siyu, et al.
Veröffentlicht: (2025)
von: Yan, Siyu, et al.
Veröffentlicht: (2025)
Tiny Refinements Elicit Resilience: Toward Efficient Prefix-Model Against LLM Red-Teaming
von: Liu, Jiaxu, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxu, et al.
Veröffentlicht: (2024)
Red Teaming Large Language Models for Healthcare
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)
FERRET: Framework for Expansion Reliant Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
von: Jang, Geonhui, et al.
Veröffentlicht: (2026)
von: Jang, Geonhui, et al.
Veröffentlicht: (2026)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
von: Yang, Yongjin, et al.
Veröffentlicht: (2024)
von: Yang, Yongjin, et al.
Veröffentlicht: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
Red Teaming Language Models for Processing Contradictory Dialogues
von: Wen, Xiaofei, et al.
Veröffentlicht: (2024)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
BioBridge: Unified Bio-Embedding with Bridging Modality in Code-Switched EMR
von: Jeon, Jangyeong, et al.
Veröffentlicht: (2024)
von: Jeon, Jangyeong, et al.
Veröffentlicht: (2024)
Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences
von: Gupta, Ayushman, et al.
Veröffentlicht: (2024)
von: Gupta, Ayushman, et al.
Veröffentlicht: (2024)
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
von: Yoo, Yongmin, et al.
Veröffentlicht: (2025)
von: Yoo, Yongmin, et al.
Veröffentlicht: (2025)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
von: Katzy, Jonathan, et al.
Veröffentlicht: (2025)
von: Katzy, Jonathan, et al.
Veröffentlicht: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
von: Yang, Yongjin, et al.
Veröffentlicht: (2024) -
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
von: Yoo, Haneul, et al.
Veröffentlicht: (2024) -
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023) -
Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
von: Yoo, Haneul, et al.
Veröffentlicht: (2025) -
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
von: Winata, Genta Indra, et al.
Veröffentlicht: (2026)