Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Guanglong, Zhang, Siyuan, Wang, Liyuan, Zhu, Jun, Su, Hang, Zhong, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
by: Lu, Keming, et al.
Published: (2024)
by: Lu, Keming, et al.
Published: (2024)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
Domain Generalizable Continual Learning
by: Yan, Hongwei, et al.
Published: (2025)
by: Yan, Hongwei, et al.
Published: (2025)
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
by: Su, Zeli, et al.
Published: (2026)
by: Su, Zeli, et al.
Published: (2026)
Noise Contrastive Alignment of Language Models with Explicit Rewards
by: Chen, Huayu, et al.
Published: (2024)
by: Chen, Huayu, et al.
Published: (2024)
Continual Learning with Global Alignment
by: Bai, Xueying, et al.
Published: (2022)
by: Bai, Xueying, et al.
Published: (2022)
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
by: Niu, Yifan, et al.
Published: (2025)
by: Niu, Yifan, et al.
Published: (2025)
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
HiDe-PET: Continual Learning via Hierarchical Decomposition of Parameter-Efficient Tuning
by: Wang, Liyuan, et al.
Published: (2024)
by: Wang, Liyuan, et al.
Published: (2024)
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Towards a General Framework for Continual Learning with Pre-training
by: Wang, Liyuan, et al.
Published: (2023)
by: Wang, Liyuan, et al.
Published: (2023)
Continual Safety Alignment via Gradient-Based Sample Selection
by: Bach, Thong, et al.
Published: (2026)
by: Bach, Thong, et al.
Published: (2026)
Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation
by: Sun, Guanglong, et al.
Published: (2025)
by: Sun, Guanglong, et al.
Published: (2025)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
by: Sun, Yitong, et al.
Published: (2026)
by: Sun, Yitong, et al.
Published: (2026)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
FlyPrompt: Brain-Inspired Random-Expanded Routing with Temporal-Ensemble Experts for General Continual Learning
by: Yan, Hongwei, et al.
Published: (2026)
by: Yan, Hongwei, et al.
Published: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026)
by: Deng, Wenlong, et al.
Published: (2026)
A Comprehensive Survey of Continual Learning: Theory, Method and Application
by: Wang, Liyuan, et al.
Published: (2023)
by: Wang, Liyuan, et al.
Published: (2023)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
by: Joshi, Maithili, et al.
Published: (2025)
by: Joshi, Maithili, et al.
Published: (2025)
Why Is RLHF Alignment Shallow? A Gradient Analysis
by: Young, Robin
Published: (2026)
by: Young, Robin
Published: (2026)
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
by: Yang, Ziyu, et al.
Published: (2026)
by: Yang, Ziyu, et al.
Published: (2026)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
by: Liu, Mingyi
Published: (2026)
by: Liu, Mingyi
Published: (2026)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
by: Li, Moxin, et al.
Published: (2025)
by: Li, Moxin, et al.
Published: (2025)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Holistic Utility Preference Learning for Listwise Alignment
by: Zhou, Jiacong, et al.
Published: (2024)
by: Zhou, Jiacong, et al.
Published: (2024)
Paying Alignment Tax with Contrastive Learning
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
by: Wang, Dianyun, et al.
Published: (2025)
by: Wang, Dianyun, et al.
Published: (2025)
MetaRM: Shifted Distributions Alignment via Meta-Learning
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
by: Janiak, Denis, et al.
Published: (2025)
by: Janiak, Denis, et al.
Published: (2025)
Course-Correction: Safety Alignment Using Synthetic Preferences
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Similar Items
-
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
by: Lu, Keming, et al.
Published: (2024) -
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023) -
Domain Generalizable Continual Learning
by: Yan, Hongwei, et al.
Published: (2025) -
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
by: Su, Zeli, et al.
Published: (2026) -
Noise Contrastive Alignment of Language Models with Explicit Rewards
by: Chen, Huayu, et al.
Published: (2024)