Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zuoou, Zhang, Weitong, Wang, Jingyuan, Zhang, Shuyuan, Bai, Wenjia, Kainz, Bernhard, Qiao, Mengyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909735318454272
author Li, Zuoou
Zhang, Weitong
Wang, Jingyuan
Zhang, Shuyuan
Bai, Wenjia
Kainz, Bernhard
Qiao, Mengyun
author_facet Li, Zuoou
Zhang, Weitong
Wang, Jingyuan
Zhang, Shuyuan
Bai, Wenjia
Kainz, Bernhard
Qiao, Mengyun
contents Multimodal large language models (MLLMs) are widely used in vision-language reasoning tasks. However, their vulnerability to adversarial prompts remains a serious concern, as safety mechanisms often fail to prevent the generation of harmful outputs. Although recent jailbreak strategies report high success rates, many responses classified as "successful" are actually benign, vague, or unrelated to the intended malicious goal. This mismatch suggests that current evaluation standards may overestimate the effectiveness of such attacks. To address this issue, we introduce a four-axis evaluation framework that considers input on-topicness, input out-of-distribution (OOD) intensity, output harmfulness, and output refusal rate. This framework identifies truly effective jailbreaks. In a substantial empirical study, we reveal a structural trade-off: highly on-topic prompts are frequently blocked by safety filters, whereas those that are too OOD often evade detection but fail to produce harmful content. However, prompts that balance relevance and novelty are more likely to evade filters and trigger dangerous output. Building on this insight, we develop a recursive rewriting strategy called Balanced Structural Decomposition (BSD). The approach restructures malicious prompts into semantically aligned sub-tasks, while introducing subtle OOD signals and visual cues that make the inputs harder to detect. BSD was tested across 13 commercial and open-source MLLMs, where it consistently led to higher attack success rates, more harmful outputs, and fewer refusals. Compared to previous methods, it improves success rates by $67\%$ and harmfulness by $21\%$, revealing a previously underappreciated weakness in current multimodal safety systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
Li, Zuoou
Zhang, Weitong
Wang, Jingyuan
Zhang, Shuyuan
Bai, Wenjia
Kainz, Bernhard
Qiao, Mengyun
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal large language models (MLLMs) are widely used in vision-language reasoning tasks. However, their vulnerability to adversarial prompts remains a serious concern, as safety mechanisms often fail to prevent the generation of harmful outputs. Although recent jailbreak strategies report high success rates, many responses classified as "successful" are actually benign, vague, or unrelated to the intended malicious goal. This mismatch suggests that current evaluation standards may overestimate the effectiveness of such attacks. To address this issue, we introduce a four-axis evaluation framework that considers input on-topicness, input out-of-distribution (OOD) intensity, output harmfulness, and output refusal rate. This framework identifies truly effective jailbreaks. In a substantial empirical study, we reveal a structural trade-off: highly on-topic prompts are frequently blocked by safety filters, whereas those that are too OOD often evade detection but fail to produce harmful content. However, prompts that balance relevance and novelty are more likely to evade filters and trigger dangerous output. Building on this insight, we develop a recursive rewriting strategy called Balanced Structural Decomposition (BSD). The approach restructures malicious prompts into semantically aligned sub-tasks, while introducing subtle OOD signals and visual cues that make the inputs harder to detect. BSD was tested across 13 commercial and open-source MLLMs, where it consistently led to higher attack success rates, more harmful outputs, and fewer refusals. Compared to previous methods, it improves success rates by $67\%$ and harmfulness by $21\%$, revealing a previously underappreciated weakness in current multimodal safety systems.
title Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.09218