JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Xi, Jian, Songlei, Li, Shasha, Li, Xiaopeng, Li, Zhaoye, Ji, Bin, Wang, Baosheng, Yu, Jie
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918274105606144
author Wang, Xi
Jian, Songlei
Li, Shasha
Li, Xiaopeng
Li, Zhaoye
Ji, Bin
Wang, Baosheng
Yu, Jie
author_facet Wang, Xi
Jian, Songlei
Li, Shasha
Li, Xiaopeng
Li, Zhaoye
Ji, Bin
Wang, Baosheng
Yu, Jie
contents Despite extensive safety alignment, Large Language Models (LLMs) often fail against jailbreak attacks. While machine unlearning has emerged as a promising defense by erasing specific harmful parameters, current methods remain vulnerable to diverse jailbreaks. We first conduct an empirical study and discover that this failure mechanism is caused by jailbreaks primarily activating non-erased parameters in the intermediate layers. Further, by probing the underlying mechanism through which these circumvented parameters reassemble into the prohibited output, we verify the persistent existence of dynamic $\textbf{jailbreak paths}$ and show that the inability to rectify them constitutes the fundamental gap in existing unlearning defenses. To bridge this gap, we propose $\textbf{J}$ailbreak $\textbf{P}$ath $\textbf{U}$nlearning (JPU), which is the first to rectify dynamic jailbreak paths towards safety anchors by dynamically mining on-policy adversarial samples to expose vulnerabilities and identify jailbreak paths. Extensive experiments demonstrate that JPU significantly enhances jailbreak resistance against dynamic attacks while preserving the model's utility.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03005
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
Wang, Xi
Jian, Songlei
Li, Shasha
Li, Xiaopeng
Li, Zhaoye
Ji, Bin
Wang, Baosheng
Yu, Jie
Cryptography and Security
Artificial Intelligence
Despite extensive safety alignment, Large Language Models (LLMs) often fail against jailbreak attacks. While machine unlearning has emerged as a promising defense by erasing specific harmful parameters, current methods remain vulnerable to diverse jailbreaks. We first conduct an empirical study and discover that this failure mechanism is caused by jailbreaks primarily activating non-erased parameters in the intermediate layers. Further, by probing the underlying mechanism through which these circumvented parameters reassemble into the prohibited output, we verify the persistent existence of dynamic $\textbf{jailbreak paths}$ and show that the inability to rectify them constitutes the fundamental gap in existing unlearning defenses. To bridge this gap, we propose $\textbf{J}$ailbreak $\textbf{P}$ath $\textbf{U}$nlearning (JPU), which is the first to rectify dynamic jailbreak paths towards safety anchors by dynamically mining on-policy adversarial samples to expose vulnerabilities and identify jailbreak paths. Extensive experiments demonstrate that JPU significantly enhances jailbreak resistance against dynamic attacks while preserving the model's utility.
title JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2601.03005