JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xi, Jian, Songlei, Li, Shasha, Li, Xiaopeng, Li, Zhaoye, Ji, Bin, Wang, Baosheng, Yu, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
by: Pu, Rui, et al.
Published: (2025)
by: Pu, Rui, et al.
Published: (2025)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025)
by: Li, Xiaohu, et al.
Published: (2025)
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
by: Lin, Zhihao, et al.
Published: (2024)
by: Lin, Zhihao, et al.
Published: (2024)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
MADE: Graph Backdoor Defense with Masked Unlearning
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
by: Li, Hongyi, et al.
Published: (2024)
by: Li, Hongyi, et al.
Published: (2024)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense
by: Liu, Cheng, et al.
Published: (2026)
by: Liu, Cheng, et al.
Published: (2026)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
by: Wang, Zhaoqi, et al.
Published: (2025)
by: Wang, Zhaoqi, et al.
Published: (2025)
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
by: Cunningham, Hoagy, et al.
Published: (2026)
by: Cunningham, Hoagy, et al.
Published: (2026)
CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
by: Hu, Jiaming, et al.
Published: (2025)
by: Hu, Jiaming, et al.
Published: (2025)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
by: Liu, Ziyao, et al.
Published: (2024)
by: Liu, Ziyao, et al.
Published: (2024)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
The Path To Autonomous Cyber Defense
by: Oesch, Sean, et al.
Published: (2024)
by: Oesch, Sean, et al.
Published: (2024)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks
by: Yu, Tianlong, et al.
Published: (2026)
by: Yu, Tianlong, et al.
Published: (2026)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
by: Long, Zhuohang, et al.
Published: (2025)
by: Long, Zhuohang, et al.
Published: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization
by: Liu, Aofan, et al.
Published: (2025)
by: Liu, Aofan, et al.
Published: (2025)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
by: Chen, Kejia, et al.
Published: (2026)
by: Chen, Kejia, et al.
Published: (2026)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
by: Nian, Yi, et al.
Published: (2025)
by: Nian, Yi, et al.
Published: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference
by: Li, Zhengyi, et al.
Published: (2026)
by: Li, Zhengyi, et al.
Published: (2026)
Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
by: Liu, Zehao, et al.
Published: (2025)
by: Liu, Zehao, et al.
Published: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
by: Chen, Luoyu, et al.
Published: (2026)
by: Chen, Luoyu, et al.
Published: (2026)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
by: Ji, Zimo, et al.
Published: (2025)
by: Ji, Zimo, et al.
Published: (2025)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
by: Hong, Wenjing, et al.
Published: (2026)
by: Hong, Wenjing, et al.
Published: (2026)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
Similar Items
-
Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
by: Wang, Xi, et al.
Published: (2025) -
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
by: Lin, Zheng, et al.
Published: (2026) -
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
by: Pu, Rui, et al.
Published: (2025) -
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025) -
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
by: Lin, Zhihao, et al.
Published: (2024)