LLM Jailbreak Detection for (Almost) Free!

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Guorui, Xia, Yifan, Jia, Xiaojun, Li, Zhijiang, Torr, Philip, Gu, Jindong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912842329882624
author Chen, Guorui
Xia, Yifan
Jia, Xiaojun
Li, Zhijiang
Torr, Philip
Gu, Jindong
author_facet Chen, Guorui
Xia, Yifan
Jia, Xiaojun
Li, Zhijiang
Torr, Philip
Gu, Jindong
contents Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitigating jailbreak attacks through the assistance of other models or multiple model inferences. However, existing methods entail significant computational costs. In this paper, we first present a finding that the difference in output distributions between jailbreak and benign prompts can be employed for detecting jailbreak prompts. Based on this finding, we propose a Free Jailbreak Detection (FJD) which prepends an affirmative instruction to the input and scales the logits by temperature to further distinguish between jailbreak and benign prompts through the confidence of the first token. Furthermore, we enhance the detection performance of FJD through the integration of virtual instruction learning. Extensive experiments on aligned LLMs show that our FJD can effectively detect jailbreak prompts with almost no additional computational costs during LLM inference.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Jailbreak Detection for (Almost) Free!
Chen, Guorui
Xia, Yifan
Jia, Xiaojun
Li, Zhijiang
Torr, Philip
Gu, Jindong
Cryptography and Security
Artificial Intelligence
Computation and Language
Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitigating jailbreak attacks through the assistance of other models or multiple model inferences. However, existing methods entail significant computational costs. In this paper, we first present a finding that the difference in output distributions between jailbreak and benign prompts can be employed for detecting jailbreak prompts. Based on this finding, we propose a Free Jailbreak Detection (FJD) which prepends an affirmative instruction to the input and scales the logits by temperature to further distinguish between jailbreak and benign prompts through the confidence of the first token. Furthermore, we enhance the detection performance of FJD through the integration of virtual instruction learning. Extensive experiments on aligned LLMs show that our FJD can effectively detect jailbreak prompts with almost no additional computational costs during LLM inference.
title LLM Jailbreak Detection for (Almost) Free!
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.14558