Re-Triggering Safeguards within LLMs for Jailbreak Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Zheng, Niu, Zhenxing, Ji, Haoxuan, Huang, Yuzhe, Gao, Haichang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916000881967104
author Lin, Zheng
Niu, Zhenxing
Ji, Haoxuan
Huang, Yuzhe
Gao, Haichang
author_facet Lin, Zheng
Niu, Zhenxing
Ji, Haoxuan
Huang, Yuzhe
Gao, Haichang
contents This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts are inherently fragile, and thus introduce an embedding disruption method to re-activate the safeguards within LLMs. Unlike previous defense methods that aim to serve as standalone solutions, our approach instead cooperates with the LLM's internal defense mechanisms by re-triggering them. Moreover, through extensive analysis, we gain a comprehensive understanding of the disruption effects and develop an efficient search algorithm to identify appropriate disruptions for effective jailbreak detection. Extensive experiments demonstrate that our approach effectively defends against state-of-the-art jailbreak attacks in white-box and black-box settings, and remains robust even against adaptive attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10611
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Re-Triggering Safeguards within LLMs for Jailbreak Detection
Lin, Zheng
Niu, Zhenxing
Ji, Haoxuan
Huang, Yuzhe
Gao, Haichang
Cryptography and Security
Artificial Intelligence
This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts are inherently fragile, and thus introduce an embedding disruption method to re-activate the safeguards within LLMs. Unlike previous defense methods that aim to serve as standalone solutions, our approach instead cooperates with the LLM's internal defense mechanisms by re-triggering them. Moreover, through extensive analysis, we gain a comprehensive understanding of the disruption effects and develop an efficient search algorithm to identify appropriate disruptions for effective jailbreak detection. Extensive experiments demonstrate that our approach effectively defends against state-of-the-art jailbreak attacks in white-box and black-box settings, and remains robust even against adaptive attacks.
title Re-Triggering Safeguards within LLMs for Jailbreak Detection
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2605.10611