Saved in:
Bibliographic Details
Main Authors: Hsiung, Lei, Pang, Tianyu, Tang, Yung-Chen, Song, Linyue, Ho, Tsung-Yi, Chen, Pin-Yu, Yang, Yaoqing
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.05346
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918046489116672
author Hsiung, Lei
Pang, Tianyu
Tang, Yung-Chen
Song, Linyue
Ho, Tsung-Yi
Chen, Pin-Yu
Yang, Yaoqing
author_facet Hsiung, Lei
Pang, Tianyu
Tang, Yung-Chen
Song, Linyue
Ho, Tsung-Yi
Chen, Pin-Yu
Yang, Yaoqing
contents Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. However, existing mitigation strategies primarily focus on reactively addressing jailbreak incidents after safety guardrails have been compromised, removing harmful gradients during fine-tuning, or continuously reinforcing safety alignment throughout fine-tuning. As such, they tend to overlook a critical upstream factor: the role of the original safety-alignment data. This paper therefore investigates the degradation of safety guardrails through the lens of representation similarity between upstream alignment datasets and downstream fine-tuning tasks. Our experiments demonstrate that high similarity between these datasets significantly weakens safety guardrails, making models more susceptible to jailbreaks. Conversely, low similarity between these two types of datasets yields substantially more robust models and thus reduces harmfulness score by up to 10.33%. By highlighting the importance of upstream dataset design in the building of durable safety guardrails and reducing real-world vulnerability to jailbreak attacks, these findings offer actionable insights for fine-tuning service providers.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05346
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
Hsiung, Lei
Pang, Tianyu
Tang, Yung-Chen
Song, Linyue
Ho, Tsung-Yi
Chen, Pin-Yu
Yang, Yaoqing
Cryptography and Security
Computation and Language
Machine Learning
Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. However, existing mitigation strategies primarily focus on reactively addressing jailbreak incidents after safety guardrails have been compromised, removing harmful gradients during fine-tuning, or continuously reinforcing safety alignment throughout fine-tuning. As such, they tend to overlook a critical upstream factor: the role of the original safety-alignment data. This paper therefore investigates the degradation of safety guardrails through the lens of representation similarity between upstream alignment datasets and downstream fine-tuning tasks. Our experiments demonstrate that high similarity between these datasets significantly weakens safety guardrails, making models more susceptible to jailbreaks. Conversely, low similarity between these two types of datasets yields substantially more robust models and thus reduces harmfulness score by up to 10.33%. By highlighting the importance of upstream dataset design in the building of durable safety guardrails and reducing real-world vulnerability to jailbreak attacks, these findings offer actionable insights for fine-tuning service providers.
title Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
topic Cryptography and Security
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.05346