JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Piet, Julien, Huang, Xiao, Jacob, Dennis, Chow, Annabella, Alrashed, Maha, Zhao, Geng, Hu, Zhanhao, Sitawarin, Chawin, Alomair, Basel, Wagner, David
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913810258853888
author Piet, Julien
Huang, Xiao
Jacob, Dennis
Chow, Annabella
Alrashed, Maha
Zhao, Geng
Hu, Zhanhao
Sitawarin, Chawin
Alomair, Basel
Wagner, David
author_facet Piet, Julien
Huang, Xiao
Jacob, Dennis
Chow, Annabella
Alrashed, Maha
Zhao, Geng
Hu, Zhanhao
Sitawarin, Chawin
Alomair, Basel
Wagner, David
contents Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulnerable to jailbreak attacks that bypass safety guardrails. Universal jailbreaks - prefixes that can circumvent alignment for any payload - are particularly concerning. We show empirically that jailbreak detection systems face distribution shift, with detectors trained at one point in time performing poorly against newer exploits. To study this problem, we release JailbreaksOverTime, a comprehensive dataset of timestamped real user interactions containing both benign requests and jailbreak attempts collected over 10 months. We propose a two-pronged method for defenders to detect new jailbreaks and continuously update their detectors. First, we show how to use continuous learning to detect jailbreaks and adapt rapidly to new emerging jailbreaks. While detectors trained at a single point in time eventually fail due to drift, we find that universal jailbreaks evolve slowly enough for self-training to be effective. Retraining our detection model weekly using its own labels - with no new human labels - reduces the false negative rate from 4% to 0.3% at a false positive rate of 0.1%. Second, we introduce an unsupervised active monitoring approach to identify novel jailbreaks. Rather than classifying inputs directly, we recognize jailbreaks by their behavior, specifically, their ability to trigger models to respond to known-harmful prompts. This approach has a higher false negative rate (4.1%) than supervised methods, but it successfully identified some out-of-distribution attacks that were missed by the continuous learning approach.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19440
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
Piet, Julien
Huang, Xiao
Jacob, Dennis
Chow, Annabella
Alrashed, Maha
Zhao, Geng
Hu, Zhanhao
Sitawarin, Chawin
Alomair, Basel
Wagner, David
Cryptography and Security
Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulnerable to jailbreak attacks that bypass safety guardrails. Universal jailbreaks - prefixes that can circumvent alignment for any payload - are particularly concerning. We show empirically that jailbreak detection systems face distribution shift, with detectors trained at one point in time performing poorly against newer exploits. To study this problem, we release JailbreaksOverTime, a comprehensive dataset of timestamped real user interactions containing both benign requests and jailbreak attempts collected over 10 months. We propose a two-pronged method for defenders to detect new jailbreaks and continuously update their detectors. First, we show how to use continuous learning to detect jailbreaks and adapt rapidly to new emerging jailbreaks. While detectors trained at a single point in time eventually fail due to drift, we find that universal jailbreaks evolve slowly enough for self-training to be effective. Retraining our detection model weekly using its own labels - with no new human labels - reduces the false negative rate from 4% to 0.3% at a false positive rate of 0.1%. Second, we introduce an unsupervised active monitoring approach to identify novel jailbreaks. Rather than classifying inputs directly, we recognize jailbreaks by their behavior, specifically, their ability to trigger models to respond to known-harmful prompts. This approach has a higher false negative rate (4.1%) than supervised methods, but it successfully identified some out-of-distribution attacks that were missed by the continuous learning approach.
title JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
topic Cryptography and Security
url https://arxiv.org/abs/2504.19440