Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xiong, Yuan, Miao, Ziqi, Li, Lijun, Qian, Chen, Li, Jie, Shao, Jing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912743865450496
author Xiong, Yuan
Miao, Ziqi
Li, Lijun
Qian, Chen
Li, Jie
Shao, Jing
author_facet Xiong, Yuan
Miao, Ziqi
Li, Lijun
Qian, Chen
Li, Jie
Shao, Jing
contents While Multimodal Large Language Models (MLLMs) show remarkable capabilities, their safety alignments are susceptible to jailbreak attacks. Existing attack methods typically focus on text-image interplay, treating the visual modality as a secondary prompt. This approach underutilizes the unique potential of images to carry complex, contextual information. To address this gap, we propose a new image-centric attack method, Contextual Image Attack (CIA), which employs a multi-agent system to subtly embeds harmful queries into seemingly benign visual contexts using four distinct visualization strategies. To further enhance the attack's efficacy, the system incorporate contextual element enhancement and automatic toxicity obfuscation techniques. Experimental results on the MMSafetyBench-tiny dataset show that CIA achieves high toxicity scores of 4.73 and 4.83 against the GPT-4o and Qwen2.5-VL-72B models, respectively, with Attack Success Rates (ASR) reaching 86.31\% and 91.07\%. Our method significantly outperforms prior work, demonstrating that the visual modality itself is a potent vector for jailbreaking advanced MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02973
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
Xiong, Yuan
Miao, Ziqi
Li, Lijun
Qian, Chen
Li, Jie
Shao, Jing
Computer Vision and Pattern Recognition
Computation and Language
Cryptography and Security
While Multimodal Large Language Models (MLLMs) show remarkable capabilities, their safety alignments are susceptible to jailbreak attacks. Existing attack methods typically focus on text-image interplay, treating the visual modality as a secondary prompt. This approach underutilizes the unique potential of images to carry complex, contextual information. To address this gap, we propose a new image-centric attack method, Contextual Image Attack (CIA), which employs a multi-agent system to subtly embeds harmful queries into seemingly benign visual contexts using four distinct visualization strategies. To further enhance the attack's efficacy, the system incorporate contextual element enhancement and automatic toxicity obfuscation techniques. Experimental results on the MMSafetyBench-tiny dataset show that CIA achieves high toxicity scores of 4.73 and 4.83 against the GPT-4o and Qwen2.5-VL-72B models, respectively, with Attack Success Rates (ASR) reaching 86.31\% and 91.07\%. Our method significantly outperforms prior work, demonstrating that the visual modality itself is a potent vector for jailbreaking advanced MLLMs.
title Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
topic Computer Vision and Pattern Recognition
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2512.02973