Metaphor-based Jailbreak Attacks on Text-to-Image Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chenyu, Wang, Lanjun, Ma, Yiwen, Li, Wenhui, Tu, Yi, Liu, An-An |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
von: Zhang, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2024)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
von: Liang, Shuang, et al.
Veröffentlicht: (2025)
von: Liang, Shuang, et al.
Veröffentlicht: (2025)
Antelope: Potent and Concealed Jailbreak Attack Strategy
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
Text is All You Need for Vision-Language Model Jailbreaking
von: Chen, Yihang, et al.
Veröffentlicht: (2026)
von: Chen, Yihang, et al.
Veröffentlicht: (2026)
Multimodal Pragmatic Jailbreak on Text-to-image Models
von: Liu, Tong, et al.
Veröffentlicht: (2024)
von: Liu, Tong, et al.
Veröffentlicht: (2024)
Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2025)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2025)
Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models
von: Jang, Sangwon, et al.
Veröffentlicht: (2025)
von: Jang, Sangwon, et al.
Veröffentlicht: (2025)
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
von: Liu, Yilian, et al.
Veröffentlicht: (2026)
von: Liu, Yilian, et al.
Veröffentlicht: (2026)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
von: Gao, Sensen, et al.
Veröffentlicht: (2024)
von: Gao, Sensen, et al.
Veröffentlicht: (2024)
Backdoor Attacks against Image-to-Image Networks
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
von: Dang, Pucheng, et al.
Veröffentlicht: (2024)
von: Dang, Pucheng, et al.
Veröffentlicht: (2024)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
von: Li, Boheng, et al.
Veröffentlicht: (2025)
von: Li, Boheng, et al.
Veröffentlicht: (2025)
VA3: Virtually Assured Amplification Attack on Probabilistic Copyright Protection for Text-to-Image Generative Models
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
INK: Inheritable Natural Backdoor Attack Against Model Distillation
von: Liu, Xiaolei, et al.
Veröffentlicht: (2023)
von: Liu, Xiaolei, et al.
Veröffentlicht: (2023)
Security Risk of Misalignment between Text and Image in Multi-modal Model
von: Wang, Xiaosen, et al.
Veröffentlicht: (2025)
von: Wang, Xiaosen, et al.
Veröffentlicht: (2025)
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
von: Chen, Junxi, et al.
Veröffentlicht: (2025)
von: Chen, Junxi, et al.
Veröffentlicht: (2025)
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
von: Yuan, Lingzhi, et al.
Veröffentlicht: (2025)
von: Yuan, Lingzhi, et al.
Veröffentlicht: (2025)
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2024)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
von: Tian, Zhihua, et al.
Veröffentlicht: (2025)
von: Tian, Zhihua, et al.
Veröffentlicht: (2025)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
von: Xu, Shuhan, et al.
Veröffentlicht: (2026)
von: Xu, Shuhan, et al.
Veröffentlicht: (2026)
CAAP: Capture-Aware Adversarial Patch Attacks on Palmprint Recognition Models
von: Liu, Renyang, et al.
Veröffentlicht: (2026)
von: Liu, Renyang, et al.
Veröffentlicht: (2026)
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
von: Liu, Jiangtao, et al.
Veröffentlicht: (2025)
von: Liu, Jiangtao, et al.
Veröffentlicht: (2025)
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
von: Chen, Moyang, et al.
Veröffentlicht: (2026)
von: Chen, Moyang, et al.
Veröffentlicht: (2026)
Superpixel Attack: Enhancing Black-box Adversarial Attack with Image-driven Division Areas
von: Oe, Issa, et al.
Veröffentlicht: (2025)
von: Oe, Issa, et al.
Veröffentlicht: (2025)
SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
CopyrightMeter: Revisiting Copyright Protection in Text-to-image Models
von: Xu, Naen, et al.
Veröffentlicht: (2024)
von: Xu, Naen, et al.
Veröffentlicht: (2024)
Region-Guided Attack on the Segment Anything Model (SAM)
von: Liu, Xiaoliang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoliang, et al.
Veröffentlicht: (2024)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
von: Meng, Xiangtao, et al.
Veröffentlicht: (2025)
von: Meng, Xiangtao, et al.
Veröffentlicht: (2025)
Undermining Image and Text Classification Algorithms Using Adversarial Attacks
von: Lunga, Langalibalele, et al.
Veröffentlicht: (2024)
von: Lunga, Langalibalele, et al.
Veröffentlicht: (2024)
RemedyGS: Defend 3D Gaussian Splatting against Computation Cost Attacks
von: Li, Yanping, et al.
Veröffentlicht: (2025)
von: Li, Yanping, et al.
Veröffentlicht: (2025)
Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image Erosion
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025) -
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
von: Zhang, Chenyu, et al.
Veröffentlicht: (2024) -
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025) -
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024) -
PLA: Prompt Learning Attack against Text-to-Image Generative Models
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)