Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters
Fuente:
arXiv
Saved in:
| Main Authors: | Mustafa, Ahmed B, Ye, Zihan, Lu, Yang, Pound, Michael P, Gowda, Shreyank N |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
by: Mustafa, Ahmed B, et al.
Published: (2025)
by: Mustafa, Ahmed B, et al.
Published: (2025)
Compression as an Adversarial Amplifier Through Decision Space Reduction
by: Evans, Lewis, et al.
Published: (2026)
by: Evans, Lewis, et al.
Published: (2026)
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
ZeroDiff++: Substantial Unseen Visual-semantic Correlation in Zero-shot Learning
by: Ye, Zihan, et al.
Published: (2026)
by: Ye, Zihan, et al.
Published: (2026)
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
by: Gowda, Shreyank N, et al.
Published: (2025)
by: Gowda, Shreyank N, et al.
Published: (2025)
FE-Adapter: Adapting Image-based Emotion Classifiers to Videos
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
Telling Stories for Common Sense Zero-Shot Action Recognition
by: Gowda, Shreyank N, et al.
Published: (2023)
by: Gowda, Shreyank N, et al.
Published: (2023)
Interpretable Zero-shot Learning with Infinite Class Concepts
by: Ye, Zihan, et al.
Published: (2025)
by: Ye, Zihan, et al.
Published: (2025)
Adversarial Robustness in Zero-Shot Learning:An Empirical Study on Class and Concept-Level Vulnerabilities
by: Peng, Zhiyuan, et al.
Published: (2025)
by: Peng, Zhiyuan, et al.
Published: (2025)
Reimagining Reality: A Comprehensive Survey of Video Inpainting Techniques
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
by: Gowda, Shreyank N, et al.
Published: (2025)
by: Gowda, Shreyank N, et al.
Published: (2025)
Continual Learning Improves Zero-Shot Action Recognition
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks
by: Zou, Quanchen, et al.
Published: (2026)
by: Zou, Quanchen, et al.
Published: (2026)
Adaptive Data Dropout: Towards Self-Regulated Learning in Deep Neural Networks
by: Gahir, Amar, et al.
Published: (2026)
by: Gahir, Amar, et al.
Published: (2026)
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?
by: Gowda, Shreyank N, et al.
Published: (2025)
by: Gowda, Shreyank N, et al.
Published: (2025)
Twin Trigger Generative Networks for Backdoor Attacks against Object Detection
by: Li, Zhiying, et al.
Published: (2024)
by: Li, Zhiying, et al.
Published: (2024)
ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning
by: Ye, Zihan, et al.
Published: (2024)
by: Ye, Zihan, et al.
Published: (2024)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
Adversarial Augmentation Training Makes Action Recognition Models More Robust to Realistic Video Distribution Shifts
by: Kim, Kiyoon, et al.
Published: (2024)
by: Kim, Kiyoon, et al.
Published: (2024)
SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning
by: Liu, Hezhao, et al.
Published: (2026)
by: Liu, Hezhao, et al.
Published: (2026)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
by: Peng, Xingkai, et al.
Published: (2025)
by: Peng, Xingkai, et al.
Published: (2025)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
by: Gao, Sensen, et al.
Published: (2024)
by: Gao, Sensen, et al.
Published: (2024)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
Watt For What: Rethinking Deep Learning's Energy-Performance Relationship
by: Gowda, Shreyank N, et al.
Published: (2023)
by: Gowda, Shreyank N, et al.
Published: (2023)
Safety-Potential Pruning for Enhancing Safety Prompts Against VLM Jailbreaking Without Retraining
by: Li, Chongxin, et al.
Published: (2026)
by: Li, Chongxin, et al.
Published: (2026)
FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled Data
by: Liu, Hezhao, et al.
Published: (2025)
by: Liu, Hezhao, et al.
Published: (2025)
CAPT: Class-Aware Prompt Tuning for Federated Long-Tailed Learning with Vision-Language Model
by: Hou, Shihao, et al.
Published: (2025)
by: Hou, Shihao, et al.
Published: (2025)
Metaphor-based Jailbreak Attacks on Text-to-Image Models
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Unified Prompt Attack Against Text-to-Image Generation Models
by: Peng, Duo, et al.
Published: (2025)
by: Peng, Duo, et al.
Published: (2025)
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
by: Khanal, Aja, et al.
Published: (2025)
by: Khanal, Aja, et al.
Published: (2025)
SIFT-Graph: Benchmarking Multimodal Defense Against Image Adversarial Attacks With Robust Feature Graph
by: He, Jingjie, et al.
Published: (2025)
by: He, Jingjie, et al.
Published: (2025)
Perception-guided Jailbreak against Text-to-Image Models
by: Huang, Yihao, et al.
Published: (2024)
by: Huang, Yihao, et al.
Published: (2024)
Principles of Visual Tokens for Efficient Video Understanding
by: Hao, Xinyue, et al.
Published: (2024)
by: Hao, Xinyue, et al.
Published: (2024)
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
by: Sathiyanarayanan, Shriram M, et al.
Published: (2025)
by: Sathiyanarayanan, Shriram M, et al.
Published: (2025)
Bridging the Projection Gap: Overcoming Projection Bias Through Parameterized Distance Learning
by: Zhang, Chong, et al.
Published: (2023)
by: Zhang, Chong, et al.
Published: (2023)
UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers
by: Peng, Duo, et al.
Published: (2024)
by: Peng, Duo, et al.
Published: (2024)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
Performance is not All You Need: Sustainability Considerations for Algorithms
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks
by: Liu, Jiayang, et al.
Published: (2025)
by: Liu, Jiayang, et al.
Published: (2025)
Similar Items
-
Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
by: Mustafa, Ahmed B, et al.
Published: (2025) -
Compression as an Adversarial Amplifier Through Decision Space Reduction
by: Evans, Lewis, et al.
Published: (2026) -
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation
by: Gowda, Shreyank N, et al.
Published: (2024) -
ZeroDiff++: Substantial Unseen Visual-semantic Correlation in Zero-shot Learning
by: Ye, Zihan, et al.
Published: (2026) -
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
by: Gowda, Shreyank N, et al.
Published: (2025)