Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Mengqi, Tian, Xinyu, Shen, Xin, Zou, Shu, Ni, Jinhong, Yang, Zhaoyuan, Li, Weikang, Li, Xuesong, Zhang, Jing
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914593329119232
author He, Mengqi
Tian, Xinyu
Shen, Xin
Zou, Shu
Ni, Jinhong
Yang, Zhaoyuan
Li, Weikang
Li, Xuesong
Zhang, Jing
author_facet He, Mengqi
Tian, Xinyu
Shen, Xin
Zou, Shu
Ni, Jinhong
Yang, Zhaoyuan
Li, Weikang
Li, Xuesong
Zhang, Jing
contents Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens during autoregressive decoding, and non-refusal tokens already carry substantial probability mass among the top-ranked candidates before attack. Motivated by this finding, we propose Untargeted Jailbreak via Entropy Maximization(UJEM)-KL, a lightweight attack that maximizes entropy at these decision tokens to flip refusal outcomes, while stabilizing the remaining low-entropy positions to preserve output quality. Across three VLMs and two safety benchmarks, UJEM-KL achieves competitive white-box attack success rates and consistently improves transferability, while remaining effective under representative defenses. Our experimental results indicate that the limited transferability primarily stems from overly constrained optimization objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10764
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
He, Mengqi
Tian, Xinyu
Shen, Xin
Zou, Shu
Ni, Jinhong
Yang, Zhaoyuan
Li, Weikang
Li, Xuesong
Zhang, Jing
Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10; I.4.9
Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens during autoregressive decoding, and non-refusal tokens already carry substantial probability mass among the top-ranked candidates before attack. Motivated by this finding, we propose Untargeted Jailbreak via Entropy Maximization(UJEM)-KL, a lightweight attack that maximizes entropy at these decision tokens to flip refusal outcomes, while stabilizing the remaining low-entropy positions to preserve output quality. Across three VLMs and two safety benchmarks, UJEM-KL achieves competitive white-box attack success rates and consistently improves transferability, while remaining effective under representative defenses. Our experimental results indicate that the limited transferability primarily stems from overly constrained optimization objectives.
title Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10; I.4.9
url https://arxiv.org/abs/2605.10764