Why Do Safety Guardrails Degrade Across Languages?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Max, Patel, Ameen, Truong, Sang T., Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Investigating Data Contamination for Pre-training Language Models
by: Jiang, Minhao, et al.
Published: (2024)
by: Jiang, Minhao, et al.
Published: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Building a Domain-specific Guardrail Model in Production
by: Niknazar, Mohammad, et al.
Published: (2024)
by: Niknazar, Mohammad, et al.
Published: (2024)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
by: Chen, Edward, et al.
Published: (2025)
by: Chen, Edward, et al.
Published: (2025)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)
by: Obbad, Elyas, et al.
Published: (2024)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
by: Ilin, Aleksei, et al.
Published: (2025)
by: Ilin, Aleksei, et al.
Published: (2025)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
by: Krishna, Kundan, et al.
Published: (2025)
by: Krishna, Kundan, et al.
Published: (2025)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
by: Patel, Fagun, et al.
Published: (2025)
by: Patel, Fagun, et al.
Published: (2025)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025)
by: Chan, Willy, et al.
Published: (2025)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
Quantifying the Importance of Data Alignment in Downstream Model Performance
by: Chawla, Krrish, et al.
Published: (2025)
by: Chawla, Krrish, et al.
Published: (2025)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
Why Larger Language Models Do In-context Learning Differently?
by: Shi, Zhenmei, et al.
Published: (2024)
by: Shi, Zhenmei, et al.
Published: (2024)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
by: Park, Jinyoung, et al.
Published: (2023)
by: Park, Jinyoung, et al.
Published: (2023)
Is Pre-training Truly Better Than Meta-Learning?
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
by: Islam, Akif, et al.
Published: (2025)
by: Islam, Akif, et al.
Published: (2025)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
by: Neill, James O', et al.
Published: (2025)
by: Neill, James O', et al.
Published: (2025)
KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
by: Mo, Belinda, et al.
Published: (2025)
by: Mo, Belinda, et al.
Published: (2025)
UQ: Assessing Language Models on Unsolved Questions
by: Nie, Fan, et al.
Published: (2025)
by: Nie, Fan, et al.
Published: (2025)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
by: Truong, Sang T., et al.
Published: (2024)
by: Truong, Sang T., et al.
Published: (2024)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
by: Chen, Yurun, et al.
Published: (2026)
by: Chen, Yurun, et al.
Published: (2026)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
Similar Items
-
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025) -
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
by: Shahani, Prithviraj Singh, et al.
Published: (2025) -
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025) -
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026) -
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025)