Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brito, Iago Alves, Rios, Walcy Santos Rezende, Dollis, Julia Soares, Silva, Diogo Fernandes Costa, Filho, Arlindo Rodrigues Galvão |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToxSyn: Reducing Bias in Hate Speech Detection via Synthetic Minority Data in Brazilian Portuguese
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025)
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025)
Integrating Personality into Digital Humans: A Review of LLM-Driven Approaches for Virtual Reality
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025)
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025)
MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese Speakers
von: Färber, Fernanda Bufon, et al.
Veröffentlicht: (2025)
von: Färber, Fernanda Bufon, et al.
Veröffentlicht: (2025)
When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical Training
von: Dollis, Julia S., et al.
Veröffentlicht: (2025)
von: Dollis, Julia S., et al.
Veröffentlicht: (2025)
Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
von: Gomes, Juliana Resplande Sant'anna, et al.
Veröffentlicht: (2025)
von: Gomes, Juliana Resplande Sant'anna, et al.
Veröffentlicht: (2025)
Tagarela - A Portuguese speech dataset from podcasts
von: de Oliveira, Frederico Santos, et al.
Veröffentlicht: (2026)
von: de Oliveira, Frederico Santos, et al.
Veröffentlicht: (2026)
Advancing LLM Safe Alignment with Safety Representation Ranking
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
Multilingual Blending: LLM Safety Alignment Evaluation with Language Mixture
von: Song, Jiayang, et al.
Veröffentlicht: (2024)
von: Song, Jiayang, et al.
Veröffentlicht: (2024)
AKCIT-FN at CheckThat! 2025: Switching Fine-Tuned SLMs and LLM Prompting for Multilingual Claim Normalization
von: Almada, Fabrycio Leite Nakano, et al.
Veröffentlicht: (2025)
von: Almada, Fabrycio Leite Nakano, et al.
Veröffentlicht: (2025)
Attention Guidance through Video Script: A Case Study of Object Focusing on 360° VR Video Tours
von: Silva, Paulo Vitor Santana, et al.
Veröffentlicht: (2026)
von: Silva, Paulo Vitor Santana, et al.
Veröffentlicht: (2026)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
von: Horal, Artur, et al.
Veröffentlicht: (2025)
von: Horal, Artur, et al.
Veröffentlicht: (2025)
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
von: Lou, Xinyue, et al.
Veröffentlicht: (2025)
von: Lou, Xinyue, et al.
Veröffentlicht: (2025)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
von: Yang, Yahan, et al.
Veröffentlicht: (2025)
von: Yang, Yahan, et al.
Veröffentlicht: (2025)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
What Matters For Safety Alignment?
von: Li, Xing, et al.
Veröffentlicht: (2026)
von: Li, Xing, et al.
Veröffentlicht: (2026)
Test-Time Safety Alignment
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
STAIR: Improving Safety Alignment with Introspective Reasoning
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2025)
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
von: Sun, Chenkai, et al.
Veröffentlicht: (2025)
von: Sun, Chenkai, et al.
Veröffentlicht: (2025)
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
von: Hao, Haochang, et al.
Veröffentlicht: (2026)
von: Hao, Haochang, et al.
Veröffentlicht: (2026)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
von: Bu, Yuyan, et al.
Veröffentlicht: (2026)
von: Bu, Yuyan, et al.
Veröffentlicht: (2026)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs
von: Zhu, Han, et al.
Veröffentlicht: (2026)
von: Zhu, Han, et al.
Veröffentlicht: (2026)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
von: Hazra, Rima, et al.
Veröffentlicht: (2024)
von: Hazra, Rima, et al.
Veröffentlicht: (2024)
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
von: Qiu, Haoyi, et al.
Veröffentlicht: (2025)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2025)
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
Cat-DPO: Category-Adaptive Safety Alignment
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ToxSyn: Reducing Bias in Hate Speech Detection via Synthetic Minority Data in Brazilian Portuguese
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025) -
Integrating Personality into Digital Humans: A Review of LLM-Driven Approaches for Virtual Reality
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025) -
MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese Speakers
von: Färber, Fernanda Bufon, et al.
Veröffentlicht: (2025) -
When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical Training
von: Dollis, Julia S., et al.
Veröffentlicht: (2025) -
Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
von: Gomes, Juliana Resplande Sant'anna, et al.
Veröffentlicht: (2025)