SDD: Self-Degraded Defense against Malicious Fine-tuning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Zixuan, Lu, Weikai, Lin, Xin, Zeng, Ziqian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912507398979584
author Chen, Zixuan
Lu, Weikai
Lin, Xin
Zeng, Ziqian
author_facet Chen, Zixuan
Lu, Weikai
Lin, Xin
Zeng, Ziqian
contents Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuning succeeds and identify potential defense strategies. Building on the theoretical analysis, we introduce the Self-Degraded Defense (SDD) framework. SDD encourages LLMs to produce high-quality but irrelevant responses to harmful prompts. When attackers attempt malicious fine-tuning, the general capability of the LLM aligned by SDD will significantly decrease, rendering it incapable of following harmful instructions. Our experimental results confirm SDD's effectiveness against such attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21182
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SDD: Self-Degraded Defense against Malicious Fine-tuning
Chen, Zixuan
Lu, Weikai
Lin, Xin
Zeng, Ziqian
Cryptography and Security
Artificial Intelligence
Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuning succeeds and identify potential defense strategies. Building on the theoretical analysis, we introduce the Self-Degraded Defense (SDD) framework. SDD encourages LLMs to produce high-quality but irrelevant responses to harmful prompts. When attackers attempt malicious fine-tuning, the general capability of the LLM aligned by SDD will significantly decrease, rendering it incapable of following harmful instructions. Our experimental results confirm SDD's effectiveness against such attacks.
title SDD: Self-Degraded Defense against Malicious Fine-tuning
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2507.21182