Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Run, Ying, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PLA: Prompt Learning Attack against Text-to-Image Generative Models
by: Lyu, Xinqi, et al.
Published: (2025)
by: Lyu, Xinqi, et al.
Published: (2025)
Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models
by: Müller, Andreas, et al.
Published: (2024)
by: Müller, Andreas, et al.
Published: (2024)
Everywhere Attack: Attacking Locally and Globally to Boost Targeted Transferability
by: Zeng, Hui, et al.
Published: (2025)
by: Zeng, Hui, et al.
Published: (2025)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
StyleFool: Fooling Video Classification Systems via Style Transfer
by: Cao, Yuxin, et al.
Published: (2022)
by: Cao, Yuxin, et al.
Published: (2022)
Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D Modeling
by: Hu, Zhanhao, et al.
Published: (2023)
by: Hu, Zhanhao, et al.
Published: (2023)
Fight Perturbations with Perturbations: Defending Adversarial Attacks via Neuron Influence
by: Chen, Ruoxi, et al.
Published: (2021)
by: Chen, Ruoxi, et al.
Published: (2021)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image Erosion
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
by: Xu, Shuhan, et al.
Published: (2026)
by: Xu, Shuhan, et al.
Published: (2026)
PAD-FT: A Lightweight Defense for Backdoor Attacks via Data Purification and Fine-Tuning
by: Xu, Yukai, et al.
Published: (2024)
by: Xu, Yukai, et al.
Published: (2024)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
by: Li, Changyue, et al.
Published: (2025)
by: Li, Changyue, et al.
Published: (2025)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
by: Lee, Junhoo, et al.
Published: (2026)
by: Lee, Junhoo, et al.
Published: (2026)
Superpixel Attack: Enhancing Black-box Adversarial Attack with Image-driven Division Areas
by: Oe, Issa, et al.
Published: (2025)
by: Oe, Issa, et al.
Published: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Data Free Backdoor Attacks
by: Cao, Bochuan, et al.
Published: (2024)
by: Cao, Bochuan, et al.
Published: (2024)
Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
Improving the Perturbation-Based Explanation of Deepfake Detectors Through the Use of Adversarially-Generated Samples
by: Tsigos, Konstantinos, et al.
Published: (2025)
by: Tsigos, Konstantinos, et al.
Published: (2025)
Certified but Fooled! Breaking Certified Defences with Ghost Certificates
by: Vo, Quoc Viet, et al.
Published: (2025)
by: Vo, Quoc Viet, et al.
Published: (2025)
Backdoor Attacks against Image-to-Image Networks
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Antelope: Potent and Concealed Jailbreak Attack Strategy
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
Metaphor-based Jailbreak Attacks on Text-to-Image Models
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Region-Guided Attack on the Segment Anything Model (SAM)
by: Liu, Xiaoliang, et al.
Published: (2024)
by: Liu, Xiaoliang, et al.
Published: (2024)
Edge-Only Universal Adversarial Attacks in Distributed Learning
by: Rossolini, Giulio, et al.
Published: (2024)
by: Rossolini, Giulio, et al.
Published: (2024)
Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
by: Maljkovic, Igor, et al.
Published: (2026)
by: Maljkovic, Igor, et al.
Published: (2026)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
by: Lu, Jialin, et al.
Published: (2025)
by: Lu, Jialin, et al.
Published: (2025)
BadViM: Backdoor Attack against Vision Mamba
by: Wu, Yinghao, et al.
Published: (2025)
by: Wu, Yinghao, et al.
Published: (2025)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
by: Lu, Jialin, et al.
Published: (2024)
by: Lu, Jialin, et al.
Published: (2024)
INK: Inheritable Natural Backdoor Attack Against Model Distillation
by: Liu, Xiaolei, et al.
Published: (2023)
by: Liu, Xiaolei, et al.
Published: (2023)
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
by: Dang, Pucheng, et al.
Published: (2024)
by: Dang, Pucheng, et al.
Published: (2024)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
by: Liang, Shuang, et al.
Published: (2025)
by: Liang, Shuang, et al.
Published: (2025)
Robustness Analysis against Adversarial Patch Attacks in Fully Unmanned Stores
by: Na, Hyunsik, et al.
Published: (2025)
by: Na, Hyunsik, et al.
Published: (2025)
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025)
by: Song, Min Geun, et al.
Published: (2025)
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
CAAP: Capture-Aware Adversarial Patch Attacks on Palmprint Recognition Models
by: Liu, Renyang, et al.
Published: (2026)
by: Liu, Renyang, et al.
Published: (2026)
A Method to Facilitate Membership Inference Attacks in Deep Learning Models
by: Chen, Zitao, et al.
Published: (2024)
by: Chen, Zitao, et al.
Published: (2024)
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
by: Zhang, Chenyu, et al.
Published: (2024)
by: Zhang, Chenyu, et al.
Published: (2024)
Similar Items
-
PLA: Prompt Learning Attack against Text-to-Image Generative Models
by: Lyu, Xinqi, et al.
Published: (2025) -
Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models
by: Müller, Andreas, et al.
Published: (2024) -
Everywhere Attack: Attacking Locally and Globally to Boost Targeted Transferability
by: Zeng, Hui, et al.
Published: (2025) -
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026) -
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
by: Zhang, Chenyu, et al.
Published: (2025)