SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
Fuente:
arXiv
Saved in:
| Main Authors: | Joshi, Maithili, Nandi, Palash, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
by: Nandi, Palash, et al.
Published: (2025)
by: Nandi, Palash, et al.
Published: (2025)
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
by: Hajra, Suvadeep, et al.
Published: (2026)
by: Hajra, Suvadeep, et al.
Published: (2026)
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
by: Nandi, Palash, et al.
Published: (2024)
by: Nandi, Palash, et al.
Published: (2024)
Uncovering Cross-Objective Interference in Multi-Objective Alignment
by: Lu, Yining, et al.
Published: (2026)
by: Lu, Yining, et al.
Published: (2026)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
by: Shetty, Anudeex, et al.
Published: (2026)
by: Shetty, Anudeex, et al.
Published: (2026)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
Persona-aware Generative Model for Code-mixed Language
by: Sengupta, Ayan, et al.
Published: (2023)
by: Sengupta, Ayan, et al.
Published: (2023)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
by: Dutta, Subhabrata, et al.
Published: (2024)
by: Dutta, Subhabrata, et al.
Published: (2024)
Temporally Consistent Factuality Probing for Large Language Models
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
by: Singh, Joykirat, et al.
Published: (2025)
by: Singh, Joykirat, et al.
Published: (2025)
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
by: Sun, Guanglong, et al.
Published: (2026)
by: Sun, Guanglong, et al.
Published: (2026)
Layer by Layer: Uncovering Hidden Representations in Language Models
by: Skean, Oscar, et al.
Published: (2025)
by: Skean, Oscar, et al.
Published: (2025)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Dissecting Language Models: Machine Unlearning via Selective Pruning
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
Multilingual Language Models Encode Script Over Linguistic Structure
by: Verma, Aastha A K, et al.
Published: (2026)
by: Verma, Aastha A K, et al.
Published: (2026)
PerSEval: Assessing Personalization in Text Summarizers
by: Dasgupta, Sourish, et al.
Published: (2024)
by: Dasgupta, Sourish, et al.
Published: (2024)
Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
by: Zhao, Zheng, et al.
Published: (2024)
by: Zhao, Zheng, et al.
Published: (2024)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
by: Tran, Thien Q., et al.
Published: (2025)
by: Tran, Thien Q., et al.
Published: (2025)
Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training
by: Li, Qingyang, et al.
Published: (2024)
by: Li, Qingyang, et al.
Published: (2024)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
by: Kartik, Kartik, et al.
Published: (2024)
by: Kartik, Kartik, et al.
Published: (2024)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
by: Xiao, Da, et al.
Published: (2025)
by: Xiao, Da, et al.
Published: (2025)
KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
by: Zhou, Wuyang, et al.
Published: (2026)
by: Zhou, Wuyang, et al.
Published: (2026)
Universal Cross-Lingual Text Classification
by: Savant, Riya, et al.
Published: (2024)
by: Savant, Riya, et al.
Published: (2024)
Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
by: Chatterjee, Parthiv, et al.
Published: (2025)
by: Chatterjee, Parthiv, et al.
Published: (2025)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
Towards Building Efficient Sentence BERT Models using Layer Pruning
by: Shelke, Anushka, et al.
Published: (2024)
by: Shelke, Anushka, et al.
Published: (2024)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Alignment is Localized: A Causal Probe into Preference Layers
by: Chaudhury, Archie
Published: (2025)
by: Chaudhury, Archie
Published: (2025)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
by: Bai, Yushi, et al.
Published: (2026)
by: Bai, Yushi, et al.
Published: (2026)
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
by: Janiak, Denis, et al.
Published: (2025)
by: Janiak, Denis, et al.
Published: (2025)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
by: Sun, Yitong, et al.
Published: (2026)
by: Sun, Yitong, et al.
Published: (2026)
Similar Items
-
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
by: Nandi, Palash, et al.
Published: (2025) -
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
by: Hajra, Suvadeep, et al.
Published: (2026) -
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
by: Nandi, Palash, et al.
Published: (2024) -
Uncovering Cross-Objective Interference in Multi-Objective Alignment
by: Lu, Yining, et al.
Published: (2026) -
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
by: Bajpai, Ashutosh, et al.
Published: (2024)