Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Ham, Seokil, Choi, Yubin, Yang, Yujin, Cho, Seungju, Kim, Younghun, Kim, Changick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Locking Down the Finetuned LLMs Safety
by: Zhu, Minjun, et al.
Published: (2024)
by: Zhu, Minjun, et al.
Published: (2024)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Enhancing Robustness in Incremental Learning with Adversarial Training
by: Cho, Seungju, et al.
Published: (2023)
by: Cho, Seungju, et al.
Published: (2023)
Should LLM Safety Be More Than Refusing Harmful Instructions?
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
Large Language Models to Diffusion Finetuning
by: Cetin, Edoardo, et al.
Published: (2025)
by: Cetin, Edoardo, et al.
Published: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
Long-tailed Adversarial Training with Self-Distillation
by: Cho, Seungju, et al.
Published: (2025)
by: Cho, Seungju, et al.
Published: (2025)
Indirect Gradient Matching for Adversarial Robust Distillation
by: Lee, Hongsin, et al.
Published: (2023)
by: Lee, Hongsin, et al.
Published: (2023)
Representation Noising: A Defence Mechanism Against Harmful Finetuning
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Time is Encoded in the Weights of Finetuned Language Models
by: Nylund, Kai, et al.
Published: (2023)
by: Nylund, Kai, et al.
Published: (2023)
TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models
by: Zhang, Chen, et al.
Published: (2024)
by: Zhang, Chen, et al.
Published: (2024)
Finetune-Informed Pretraining Boosts Downstream Performance
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
by: Park, Byeongjun, et al.
Published: (2024)
by: Park, Byeongjun, et al.
Published: (2024)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
RICoTA: Red-teaming of In-the-wild Conversation with Test Attempts
by: Choi, Eujeong, et al.
Published: (2025)
by: Choi, Eujeong, et al.
Published: (2025)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
The Hallucination Tax of Reinforcement Finetuning
by: Song, Linxin, et al.
Published: (2025)
by: Song, Linxin, et al.
Published: (2025)
Detecting AI-Generated Videos with Spiking Neural Networks
by: Jang, Minsuk, et al.
Published: (2026)
by: Jang, Minsuk, et al.
Published: (2026)
Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMs
by: Tan, Bowen, et al.
Published: (2025)
by: Tan, Bowen, et al.
Published: (2025)
Order Independence With Finetuning
by: Brown, Katrina, et al.
Published: (2025)
by: Brown, Katrina, et al.
Published: (2025)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Representation Bending for Large Language Model Safety
by: Yousefpour, Ashkan, et al.
Published: (2025)
by: Yousefpour, Ashkan, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Improving Sparse Memory Finetuning
by: Goyal, Satyam, et al.
Published: (2026)
by: Goyal, Satyam, et al.
Published: (2026)
Predicting Emergent Capabilities by Finetuning
by: Snell, Charlie, et al.
Published: (2024)
by: Snell, Charlie, et al.
Published: (2024)
Whisper Finetuning on Nepali Language
by: Rijal, Sanjay, et al.
Published: (2024)
by: Rijal, Sanjay, et al.
Published: (2024)
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
by: Franzmeyer, Tim, et al.
Published: (2025)
by: Franzmeyer, Tim, et al.
Published: (2025)
GradShield: Alignment Preserving Finetuning
by: Hu, Zhanhao, et al.
Published: (2026)
by: Hu, Zhanhao, et al.
Published: (2026)
Understanding the Effects of Domain Finetuning on LLMs
by: Tanwar, Eshaan, et al.
Published: (2025)
by: Tanwar, Eshaan, et al.
Published: (2025)
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning
by: Cristofano, Tony
Published: (2026)
by: Cristofano, Tony
Published: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
KnowLA: Enhancing Parameter-efficient Finetuning with Knowledgeable Adaptation
by: Luo, Xindi, et al.
Published: (2024)
by: Luo, Xindi, et al.
Published: (2024)
Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning
by: Thakrar, Karishma, et al.
Published: (2025)
by: Thakrar, Karishma, et al.
Published: (2025)
Similar Items
-
Locking Down the Finetuned LLMs Safety
by: Zhu, Minjun, et al.
Published: (2024) -
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
by: Ham, Seokil, et al.
Published: (2024) -
Enhancing Robustness in Incremental Learning with Adversarial Training
by: Cho, Seungju, et al.
Published: (2023) -
Should LLM Safety Be More Than Refusing Harmful Instructions?
by: Maskey, Utsav, et al.
Published: (2025) -
Large Language Models to Diffusion Finetuning
by: Cetin, Edoardo, et al.
Published: (2025)