Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Wenxuan, Gao, Kuiyi, Yuan, Youliang, Huang, Jen-tse, Liu, Qiuzhi, Wang, Shuai, Jiao, Wenxiang, Tu, Zhaopeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
by: Yan, Song, et al.
Published: (2025)
by: Yan, Song, et al.
Published: (2025)
A Multi-task Adversarial Attack Against Face Authentication
by: Wang, Hanrui, et al.
Published: (2024)
by: Wang, Hanrui, et al.
Published: (2024)
VVRec: Reconstruction Attacks on DL-based Volumetric Video Upstreaming via Latent Diffusion Model with Gamma Distribution
by: Lu, Rui, et al.
Published: (2025)
by: Lu, Rui, et al.
Published: (2025)
From Attack to Protection: Leveraging Watermarking Attack Network for Advanced Add-on Watermarking
by: Nam, Seung-Hun, et al.
Published: (2020)
by: Nam, Seung-Hun, et al.
Published: (2020)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
VA3: Virtually Assured Amplification Attack on Probabilistic Copyright Protection for Text-to-Image Generative Models
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
Protocol as Poetry: A Case Study of Pak's Smart Contract-Based Protocol Art
by: Hu, Botao Amber
Published: (2025)
by: Hu, Botao Amber
Published: (2025)
Game mechanics for cyber-harm awareness in the metaverse
by: McKenzie, Sophie, et al.
Published: (2025)
by: McKenzie, Sophie, et al.
Published: (2025)
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
by: Lu, Weikai, et al.
Published: (2025)
by: Lu, Weikai, et al.
Published: (2025)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
Natural Language Induced Adversarial Images
by: Zhu, Xiaopei, et al.
Published: (2024)
by: Zhu, Xiaopei, et al.
Published: (2024)
BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
by: Zhang, Zheng, et al.
Published: (2024)
by: Zhang, Zheng, et al.
Published: (2024)
Blind Deep-Learning-Based Image Watermarking Robust Against Geometric Transformations
by: Mareen, Hannes, et al.
Published: (2024)
by: Mareen, Hannes, et al.
Published: (2024)
Provably Secure Robust Image Steganography via Cross-Modal Error Correction
by: Qi, Yuang, et al.
Published: (2024)
by: Qi, Yuang, et al.
Published: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models
by: Wen, Jinming, et al.
Published: (2025)
by: Wen, Jinming, et al.
Published: (2025)
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
by: Liu, Wenyang, et al.
Published: (2024)
by: Liu, Wenyang, et al.
Published: (2024)
QMedShield: A Novel Quantum Chaos-based Image Encryption Scheme for Secure Medical Image Storage in the Cloud
by: Rajan, Arun Amaithi, et al.
Published: (2024)
by: Rajan, Arun Amaithi, et al.
Published: (2024)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Activation-Guided Local Editing for Jailbreaking Attacks
by: Wang, Jiecong, et al.
Published: (2025)
by: Wang, Jiecong, et al.
Published: (2025)
Security Analysis of Thumbnail-Preserving Image Encryption and a New Framework
by: Xie, Dong, et al.
Published: (2025)
by: Xie, Dong, et al.
Published: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing
by: Qi, Leyi, et al.
Published: (2026)
by: Qi, Leyi, et al.
Published: (2026)
MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models
by: Mao, Po-Yuan, et al.
Published: (2025)
by: Mao, Po-Yuan, et al.
Published: (2025)
SyncGuard: Robust Audio Watermarking Capable of Countering Desynchronization Attacks
by: Gan, Zhenliang, et al.
Published: (2025)
by: Gan, Zhenliang, et al.
Published: (2025)
Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning
by: Liu, Xinwei, et al.
Published: (2024)
by: Liu, Xinwei, et al.
Published: (2024)
Test-Time Backdoor Attacks on Multimodal Large Language Models
by: Lu, Dong, et al.
Published: (2024)
by: Lu, Dong, et al.
Published: (2024)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Editing Away the Evidence: Diffusion-Based Image Manipulation and the Failure Modes of Robust Watermarking
by: Qi, Qian, et al.
Published: (2026)
by: Qi, Qian, et al.
Published: (2026)
PPVF: An Efficient Privacy-Preserving Online Video Fetching Framework with Correlated Differential Privacy
by: Zhang, Xianzhi, et al.
Published: (2024)
by: Zhang, Xianzhi, et al.
Published: (2024)
DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
by: Zhao, Kangran, et al.
Published: (2025)
by: Zhao, Kangran, et al.
Published: (2025)
VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2026)
by: Cao, Yuxin, et al.
Published: (2026)
CoreMark: Toward Robust and Universal Text Watermarking Technique
by: Meng, Jiale, et al.
Published: (2025)
by: Meng, Jiale, et al.
Published: (2025)
Steganography -- coding and intercepting the information from encoded pictures in the absence of any initial information
by: Kwiatkowska, Monika, et al.
Published: (2014)
by: Kwiatkowska, Monika, et al.
Published: (2014)
Wallcamera: Reinventing the Wheel?
by: Bourquard, Aurélien, et al.
Published: (2024)
by: Bourquard, Aurélien, et al.
Published: (2024)
SyncGait: Robust Long-Distance Authentication for Drone Delivery via Implicit Gait Behaviors
by: Ling, Zijian, et al.
Published: (2025)
by: Ling, Zijian, et al.
Published: (2025)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
by: Lu, Weikai, et al.
Published: (2025)
by: Lu, Weikai, et al.
Published: (2025)
Similar Items
-
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025) -
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
by: Huang, Xijie, et al.
Published: (2024) -
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
by: Yan, Song, et al.
Published: (2025) -
A Multi-task Adversarial Attack Against Face Authentication
by: Wang, Hanrui, et al.
Published: (2024) -
VVRec: Reconstruction Attacks on DL-based Volumetric Video Upstreaming via Latent Diffusion Model with Gamma Distribution
by: Lu, Rui, et al.
Published: (2025)