Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Shuang, Xu, Zhihao, Tao, Jialing, Xue, Hui, Wang, Xiting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917026329526272
author Liang, Shuang
Xu, Zhihao
Tao, Jialing
Xue, Hui
Wang, Xiting
author_facet Liang, Shuang
Xu, Zhihao
Tao, Jialing
Xue, Hui
Wang, Xiting
contents Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks, posing serious safety risks. To address this, existing detection methods either learn attack-specific parameters, which hinders generalization to unseen attacks, or rely on heuristically sound principles, which limit accuracy and efficiency. To overcome these limitations, we propose Learning to Detect (LoD), a general framework that accurately detects unknown jailbreak attacks by shifting the focus from attack-specific learning to task-specific learning. This framework includes a Multi-modal Safety Concept Activation Vector module for safety-oriented representation learning and a Safety Pattern Auto-Encoder module for unsupervised attack classification. Extensive experiments show that our method achieves consistently higher detection AUROC on diverse unknown attacks while improving efficiency. The code is available at https://anonymous.4open.science/r/Learning-to-Detect-51CB.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
Liang, Shuang
Xu, Zhihao
Tao, Jialing
Xue, Hui
Wang, Xiting
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks, posing serious safety risks. To address this, existing detection methods either learn attack-specific parameters, which hinders generalization to unseen attacks, or rely on heuristically sound principles, which limit accuracy and efficiency. To overcome these limitations, we propose Learning to Detect (LoD), a general framework that accurately detects unknown jailbreak attacks by shifting the focus from attack-specific learning to task-specific learning. This framework includes a Multi-modal Safety Concept Activation Vector module for safety-oriented representation learning and a Safety Pattern Auto-Encoder module for unsupervised attack classification. Extensive experiments show that our method achieves consistently higher detection AUROC on diverse unknown attacks while improving efficiency. The code is available at https://anonymous.4open.science/r/Learning-to-Detect-51CB.
title Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.15430