Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Lei, Zhang, Zixun, Wang, Zizhou, Sun, Xiaobing, Li, Zhen, Zhen, Liangli, Xu, Xiaohua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913904148348928
author Jiang, Lei
Zhang, Zixun
Wang, Zizhou
Sun, Xiaobing
Li, Zhen
Zhen, Liangli
Xu, Xiaohua
author_facet Jiang, Lei
Zhang, Zixun
Wang, Zizhou
Sun, Xiaobing
Li, Zhen
Zhen, Liangli
Xu, Xiaohua
contents Large Vision-Language Models (LVLMs) demonstrate exceptional performance across multimodal tasks, yet remain vulnerable to jailbreak attacks that bypass built-in safety mechanisms to elicit restricted content generation. Existing black-box jailbreak methods primarily rely on adversarial textual prompts or image perturbations, yet these approaches are highly detectable by standard content filtering systems and exhibit low query and computational efficiency. In this work, we present Cross-modal Adversarial Multimodal Obfuscation (CAMO), a novel black-box jailbreak attack framework that decomposes malicious prompts into semantically benign visual and textual fragments. By leveraging LVLMs' cross-modal reasoning abilities, CAMO covertly reconstructs harmful instructions through multi-step reasoning, evading conventional detection mechanisms. Our approach supports adjustable reasoning complexity and requires significantly fewer queries than prior attacks, enabling both stealth and efficiency. Comprehensive evaluations conducted on leading LVLMs validate CAMO's effectiveness, showcasing robust performance and strong cross-model transferability. These results underscore significant vulnerabilities in current built-in safety mechanisms, emphasizing an urgent need for advanced, alignment-aware security and safety solutions in vision-language systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
Jiang, Lei
Zhang, Zixun
Wang, Zizhou
Sun, Xiaobing
Li, Zhen
Zhen, Liangli
Xu, Xiaohua
Computation and Language
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) demonstrate exceptional performance across multimodal tasks, yet remain vulnerable to jailbreak attacks that bypass built-in safety mechanisms to elicit restricted content generation. Existing black-box jailbreak methods primarily rely on adversarial textual prompts or image perturbations, yet these approaches are highly detectable by standard content filtering systems and exhibit low query and computational efficiency. In this work, we present Cross-modal Adversarial Multimodal Obfuscation (CAMO), a novel black-box jailbreak attack framework that decomposes malicious prompts into semantically benign visual and textual fragments. By leveraging LVLMs' cross-modal reasoning abilities, CAMO covertly reconstructs harmful instructions through multi-step reasoning, evading conventional detection mechanisms. Our approach supports adjustable reasoning complexity and requires significantly fewer queries than prior attacks, enabling both stealth and efficiency. Comprehensive evaluations conducted on leading LVLMs validate CAMO's effectiveness, showcasing robust performance and strong cross-model transferability. These results underscore significant vulnerabilities in current built-in safety mechanisms, emphasizing an urgent need for advanced, alignment-aware security and safety solutions in vision-language systems.
title Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.16760