Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Goel, Jyotin, Maji, Souvik, Mazumder, Pratik |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
by: Liu, Zeyu, et al.
Published: (2024)
by: Liu, Zeyu, et al.
Published: (2024)
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
by: Xia, Yuchen, et al.
Published: (2024)
by: Xia, Yuchen, et al.
Published: (2024)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026)
by: Vennemeyer, Daniel, et al.
Published: (2026)
Semantic-Anchored, Class Variance-Optimized Clustering for Robust Semi-Supervised Few-Shot Learning
by: Maji, Souvik, et al.
Published: (2025)
by: Maji, Souvik, et al.
Published: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
by: Goel, Raghavv, et al.
Published: (2024)
by: Goel, Raghavv, et al.
Published: (2024)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
by: Halfon, Alon, et al.
Published: (2024)
by: Halfon, Alon, et al.
Published: (2024)
In-Context Learning and Fine-Tuning GPT for Argument Mining
by: Cabessa, Jérémie, et al.
Published: (2024)
by: Cabessa, Jérémie, et al.
Published: (2024)
Domain-Adaptive Small Language Models for Structured Tax Code Prediction
by: Nath, Souvik, et al.
Published: (2025)
by: Nath, Souvik, et al.
Published: (2025)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
by: Ponkshe, Kaustubh, et al.
Published: (2025)
by: Ponkshe, Kaustubh, et al.
Published: (2025)
Anchored Supervised Fine-Tuning
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
by: Perin, Gabriel J., et al.
Published: (2025)
by: Perin, Gabriel J., et al.
Published: (2025)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Prior-Informed Zeroth-Order Optimization with Adaptive Direction Alignment for Memory-Efficient LLM Fine-Tuning
by: Jin, Feihu, et al.
Published: (2026)
by: Jin, Feihu, et al.
Published: (2026)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
by: Diao, Muxi, et al.
Published: (2026)
by: Diao, Muxi, et al.
Published: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
by: Jang, Jisoo, et al.
Published: (2025)
by: Jang, Jisoo, et al.
Published: (2025)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
by: Feng, Weitao, et al.
Published: (2025)
by: Feng, Weitao, et al.
Published: (2025)
Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
Supervised Fine-Tuning as Inverse Reinforcement Learning
by: Sun, Hao
Published: (2024)
by: Sun, Hao
Published: (2024)
Regularization Through Reasoning: Systematic Improvements in Language Model Classification via Explanation-Enhanced Fine-Tuning
by: Shah, Vivswan, et al.
Published: (2025)
by: Shah, Vivswan, et al.
Published: (2025)
Fine Tuning Methods for Low-resource Languages
by: Bakkenes, Tim, et al.
Published: (2025)
by: Bakkenes, Tim, et al.
Published: (2025)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
by: Lee, Changhun, et al.
Published: (2024)
by: Lee, Changhun, et al.
Published: (2024)
Energy and Carbon Considerations of Fine-Tuning BERT
by: Wang, Xiaorong, et al.
Published: (2023)
by: Wang, Xiaorong, et al.
Published: (2023)
UFT: Unifying Supervised and Reinforcement Fine-Tuning
by: Liu, Mingyang, et al.
Published: (2025)
by: Liu, Mingyang, et al.
Published: (2025)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
by: Yano, Kazuki, et al.
Published: (2026)
by: Yano, Kazuki, et al.
Published: (2026)
Deep Learning for Medical Text Processing: BERT Model Fine-Tuning and Comparative Study
by: Hu, Jiacheng, et al.
Published: (2024)
by: Hu, Jiacheng, et al.
Published: (2024)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
by: She, Jingyuan Selena, et al.
Published: (2023)
by: She, Jingyuan Selena, et al.
Published: (2023)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
Teaching LLMs How to Learn with Contextual Fine-Tuning
by: Choi, Younwoo, et al.
Published: (2025)
by: Choi, Younwoo, et al.
Published: (2025)
Lifelong and Continual Learning Dialogue Systems
by: Mazumder, Sahisnu, et al.
Published: (2022)
by: Mazumder, Sahisnu, et al.
Published: (2022)
LIFT: Last-Mile Fine-Tuning for Table Explicitation
by: Khaitan, Divij, et al.
Published: (2026)
by: Khaitan, Divij, et al.
Published: (2026)
Causal Fine-Tuning under Latent Confounded Shift
by: Yu, Jialin, et al.
Published: (2024)
by: Yu, Jialin, et al.
Published: (2024)
Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
by: Abdalla, M. H. I., et al.
Published: (2025)
by: Abdalla, M. H. I., et al.
Published: (2025)
Parameter-Efficient Fine-Tuning via Circular Convolution
by: Chen, Aochuan, et al.
Published: (2024)
by: Chen, Aochuan, et al.
Published: (2024)
Similar Items
-
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025) -
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
by: Wang, Hao, et al.
Published: (2026) -
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
by: Liu, Zeyu, et al.
Published: (2024) -
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
by: Azizi, Seyedarmin, et al.
Published: (2024) -
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
by: Xia, Yuchen, et al.
Published: (2024)