Saved in:
Bibliographic Details
Main Authors: Hou, Jiacheng, Sun, Yining, Jin, Ruochong, Han, Haochen, Liu, Fangming, Chan, Wai Kin Victor, Wang, Alex Jinpeng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.10179
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917266385272832
author Hou, Jiacheng
Sun, Yining
Jin, Ruochong
Han, Haochen
Liu, Fangming
Chan, Wai Kin Victor
Wang, Alex Jinpeng
author_facet Hou, Jiacheng
Sun, Yining
Jin, Ruochong
Han, Haochen
Liu, Fangming
Chan, Wai Kin Victor
Wang, Alex Jinpeng
contents Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts. While this paradigm greatly expands usability, it also introduces a critical and underexplored safety risk: the attack surface itself becomes visual. In this work, we propose Vision-Centric Jailbreak Attack (VJA), the first visual-to-visual jailbreak attack that conveys malicious instructions purely through visual inputs. To systematically study this emerging threat, we introduce IESBench, a safety-oriented benchmark for image editing models. Extensive experiments on IESBench demonstrate that VJA effectively compromises state-of-the-art commercial models, achieving attack success rates of up to 80.9% on Nano Banana Pro and 70.1% on GPT-Image-1.5. To mitigate this vulnerability, we propose a training-free defense based on introspective multimodal reasoning, which substantially improves the safety of poorly aligned models to a level comparable with commercial systems, without auxiliary guard models and with negligible computational overhead. Our findings expose new vulnerabilities, provide both a benchmark and practical defense to advance safe and trustworthy modern image editing systems. Warning: This paper contains offensive images created by large image editing models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10179
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Hou, Jiacheng
Sun, Yining
Jin, Ruochong
Han, Haochen
Liu, Fangming
Chan, Wai Kin Victor
Wang, Alex Jinpeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts. While this paradigm greatly expands usability, it also introduces a critical and underexplored safety risk: the attack surface itself becomes visual. In this work, we propose Vision-Centric Jailbreak Attack (VJA), the first visual-to-visual jailbreak attack that conveys malicious instructions purely through visual inputs. To systematically study this emerging threat, we introduce IESBench, a safety-oriented benchmark for image editing models. Extensive experiments on IESBench demonstrate that VJA effectively compromises state-of-the-art commercial models, achieving attack success rates of up to 80.9% on Nano Banana Pro and 70.1% on GPT-Image-1.5. To mitigate this vulnerability, we propose a training-free defense based on introspective multimodal reasoning, which substantially improves the safety of poorly aligned models to a level comparable with commercial systems, without auxiliary guard models and with negligible computational overhead. Our findings expose new vulnerabilities, provide both a benchmark and practical defense to advance safe and trustworthy modern image editing systems. Warning: This paper contains offensive images created by large image editing models.
title When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.10179