Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinfeng, Pang, Shengyuan, Wu, Jialin, Deng, Jiangyi, Zhong, Huanlong, Chen, Yanjiao, Zhang, Jie, Xu, Wenyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
by: Li, Xinfeng, et al.
Published: (2024)
by: Li, Xinfeng, et al.
Published: (2024)
Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)
by: Zhong, Yinan, et al.
Published: (2025)
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
by: Yan, Song, et al.
Published: (2025)
by: Yan, Song, et al.
Published: (2025)
Signal Adversarial Examples Generation for Signal Detection Network via White-Box Attack
by: Li, Dongyang, et al.
Published: (2024)
by: Li, Dongyang, et al.
Published: (2024)
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
An Effective and Resilient Backdoor Attack Framework against Deep Neural Networks and Vision Transformers
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
TPatch: A Triggered Physical Adversarial Patch
by: Zhu, Wenjun, et al.
Published: (2023)
by: Zhu, Wenjun, et al.
Published: (2023)
Megatron: Evasive Clean-Label Backdoor Attacks against Vision Transformer
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
Anomaly Unveiled: Securing Image Classification against Adversarial Patch Attacks
by: Chattopadhyay, Nandish, et al.
Published: (2024)
by: Chattopadhyay, Nandish, et al.
Published: (2024)
Robust Image Classification: Defensive Strategies against FGSM and PGD Adversarial Attacks
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer
by: Deng, Jiangyi, et al.
Published: (2024)
by: Deng, Jiangyi, et al.
Published: (2024)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
SemiAdv: Query-Efficient Black-Box Adversarial Attack with Unlabeled Images
by: Fan, Mingyuan, et al.
Published: (2024)
by: Fan, Mingyuan, et al.
Published: (2024)
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
Safeguarding Medical Image Segmentation Datasets against Unauthorized Training via Contour- and Texture-Aware Perturbations
by: Lin, Xun, et al.
Published: (2024)
by: Lin, Xun, et al.
Published: (2024)
PINA: Prompt Injection Attack against Navigation Agents
by: Liu, Jiani, et al.
Published: (2026)
by: Liu, Jiani, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models
by: Wang, Yanting, et al.
Published: (2024)
by: Wang, Yanting, et al.
Published: (2024)
SAP-DIFF: Semantic Adversarial Patch Generation for Black-Box Face Recognition Models via Diffusion Models
by: Wang, Mingsi, et al.
Published: (2025)
by: Wang, Mingsi, et al.
Published: (2025)
A Random Ensemble of Encrypted models for Enhancing Robustness against Adversarial Examples
by: Iijima, Ryota, et al.
Published: (2024)
by: Iijima, Ryota, et al.
Published: (2024)
Towards Physically Realizable Adversarial Attenuation Patch against SAR Object Detection
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
ViTGuard: Attention-aware Detection against Adversarial Examples for Vision Transformer
by: Sun, Shihua, et al.
Published: (2024)
by: Sun, Shihua, et al.
Published: (2024)
Adversarial Training against Location-Optimized Adversarial Patches
by: Rao, Sukrut, et al.
Published: (2020)
by: Rao, Sukrut, et al.
Published: (2020)
Controllable Adversarial Makeup for Privacy via Text-Guided Diffusion
by: Kwon, Youngjin, et al.
Published: (2025)
by: Kwon, Youngjin, et al.
Published: (2025)
Natural Language Induced Adversarial Images
by: Zhu, Xiaopei, et al.
Published: (2024)
by: Zhu, Xiaopei, et al.
Published: (2024)
Physical 3D Adversarial Attacks against Monocular Depth Estimation in Autonomous Driving
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
FT-Shield: A Watermark Against Unauthorized Fine-tuning in Text-to-Image Diffusion Models
by: Cui, Yingqian, et al.
Published: (2023)
by: Cui, Yingqian, et al.
Published: (2023)
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
by: Zhang, Chenyu, et al.
Published: (2024)
by: Zhang, Chenyu, et al.
Published: (2024)
Agentic Copyright Watermarking against Adversarial Evidence Forgery with Purification-Agnostic Curriculum Proxy Learning
by: Bao, Erjin, et al.
Published: (2024)
by: Bao, Erjin, et al.
Published: (2024)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
by: Ba, Zhongjie, et al.
Published: (2023)
by: Ba, Zhongjie, et al.
Published: (2023)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
by: Lyu, Xinqi, et al.
Published: (2025)
by: Lyu, Xinqi, et al.
Published: (2025)
Data-free Defense of Black Box Models Against Adversarial Attacks
by: Nayak, Gaurav Kumar, et al.
Published: (2022)
by: Nayak, Gaurav Kumar, et al.
Published: (2022)
Invisible Optical Adversarial Stripes on Traffic Sign against Autonomous Vehicles
by: Guo, Dongfang, et al.
Published: (2024)
by: Guo, Dongfang, et al.
Published: (2024)
AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models
by: Wu, Zongyu, et al.
Published: (2025)
by: Wu, Zongyu, et al.
Published: (2025)
Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches
by: Pintor, Maura, et al.
Published: (2022)
by: Pintor, Maura, et al.
Published: (2022)
Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
Similar Items
-
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
by: Li, Xinfeng, et al.
Published: (2024) -
Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities
by: Wu, Jialin, et al.
Published: (2025) -
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025) -
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
by: Yan, Song, et al.
Published: (2025) -
Signal Adversarial Examples Generation for Signal Detection Network via White-Box Attack
by: Li, Dongyang, et al.
Published: (2024)