LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Guanghao, Qiu, Panjia, Chen, Cen, Li, Hongyu, Chu, Mingyuan, Zhang, Xin, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
Evaluating Psychological Safety of Large Language Models
von: Li, Xingxuan, et al.
Veröffentlicht: (2022)
von: Li, Xingxuan, et al.
Veröffentlicht: (2022)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
von: Ziheng, Zhou, et al.
Veröffentlicht: (2026)
von: Ziheng, Zhou, et al.
Veröffentlicht: (2026)
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Defining and Evaluating Physical Safety for Large Language Models
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2024)
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2024)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
von: Wang, Dianyun, et al.
Veröffentlicht: (2025)
von: Wang, Dianyun, et al.
Veröffentlicht: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
von: Nakao, Mahiro, et al.
Veröffentlicht: (2026)
von: Nakao, Mahiro, et al.
Veröffentlicht: (2026)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
von: Meadows, Gwenyth Isobel, et al.
Veröffentlicht: (2024)
von: Meadows, Gwenyth Isobel, et al.
Veröffentlicht: (2024)
LLM Safety Alignment is Divergence Estimation in Disguise
von: Haldar, Rajdeep, et al.
Veröffentlicht: (2025)
von: Haldar, Rajdeep, et al.
Veröffentlicht: (2025)
Phare: A Safety Probe for Large Language Models
von: Jeune, Pierre Le, et al.
Veröffentlicht: (2025)
von: Jeune, Pierre Le, et al.
Veröffentlicht: (2025)
Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
SafeRBench: Dissecting the Reasoning Safety of Large Language Models
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
von: Brophy, Matthew
Veröffentlicht: (2025)
von: Brophy, Matthew
Veröffentlicht: (2025)
On Fairness of Low-Rank Adaptation of Large Models
von: Ding, Zhoujie, et al.
Veröffentlicht: (2024)
von: Ding, Zhoujie, et al.
Veröffentlicht: (2024)
Case-based Reasoning Augmented Large Language Model Framework for Decision Making in Realistic Safety-Critical Driving Scenarios
von: Gan, Wenbin, et al.
Veröffentlicht: (2025)
von: Gan, Wenbin, et al.
Veröffentlicht: (2025)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
Alignment as Iatrogenesis: Pastoral Power, Collective Pathology, and the Structural Limits of Monolingual Safety Evaluation
von: Fukui, Hiroki
Veröffentlicht: (2026)
von: Fukui, Hiroki
Veröffentlicht: (2026)
Understanding Public Safety Trends in Calgary through data mining
von: Dewis, Zack, et al.
Veröffentlicht: (2024)
von: Dewis, Zack, et al.
Veröffentlicht: (2024)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
von: Dobbe, Roel
Veröffentlicht: (2025)
von: Dobbe, Roel
Veröffentlicht: (2025)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
von: Duan, Ranjie, et al.
Veröffentlicht: (2025)
von: Duan, Ranjie, et al.
Veröffentlicht: (2025)
The Violation State: Safety State Persistence in a Multimodal Language Model Interface
von: DeVilling, Bentley
Veröffentlicht: (2025)
von: DeVilling, Bentley
Veröffentlicht: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
von: Clymer, Joshua, et al.
Veröffentlicht: (2024)
von: Clymer, Joshua, et al.
Veröffentlicht: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2025)
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2025)
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
Urban Safety Perception Through the Lens of Large Multimodal Models: A Persona-based Approach
von: Beneduce, Ciro, et al.
Veröffentlicht: (2025)
von: Beneduce, Ciro, et al.
Veröffentlicht: (2025)
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
von: Lu, Haoran, et al.
Veröffentlicht: (2025)
von: Lu, Haoran, et al.
Veröffentlicht: (2025)
LLM Safety for Children
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
Towards Context-Invariant Safety Alignment for Large Language Models
von: Wang, Yixu, et al.
Veröffentlicht: (2026)
von: Wang, Yixu, et al.
Veröffentlicht: (2026)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
von: Penny-Dimri, Jahan C., et al.
Veröffentlicht: (2025)
von: Penny-Dimri, Jahan C., et al.
Veröffentlicht: (2025)
Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics
von: Baum, Kevin
Veröffentlicht: (2025)
von: Baum, Kevin
Veröffentlicht: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety
von: Du, Y.
Veröffentlicht: (2025)
von: Du, Y.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025) -
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025) -
Evaluating Psychological Safety of Large Language Models
von: Li, Xingxuan, et al.
Veröffentlicht: (2022) -
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
von: Anwar, Usman, et al.
Veröffentlicht: (2024) -
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
von: Ziheng, Zhou, et al.
Veröffentlicht: (2026)