Buster: Implanting Semantic Backdoor into Text Encoder to Mitigate NSFW Content Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xin, Chen, Xiaojun, Xuan, Yuexin, Zhao, Zhendong, Jia, Xiaojun, Li, Xinfeng, Wang, Xiaofeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures
by: Liu, Zeyao, et al.
Published: (2026)
by: Liu, Zeyao, et al.
Published: (2026)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024)
by: Xun, Yuan, et al.
Published: (2024)
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
CipherDM: Secure Three-Party Inference for Diffusion Model Sampling
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
by: Zhao, Xin, et al.
Published: (2025)
by: Zhao, Xin, et al.
Published: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Antelope: Potent and Concealed Jailbreak Attack Strategy
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
by: Yousaf, Adeel, et al.
Published: (2025)
by: Yousaf, Adeel, et al.
Published: (2025)
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
by: Li, Xinfeng, et al.
Published: (2024)
by: Li, Xinfeng, et al.
Published: (2024)
VModA: An Effective Framework for Adaptive NSFW Image Moderation
by: Bao, Han, et al.
Published: (2025)
by: Bao, Han, et al.
Published: (2025)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-form Layout-to-Image Generation
by: Jia, Chengyou, et al.
Published: (2023)
by: Jia, Chengyou, et al.
Published: (2023)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025)
by: Wang, Ruotong, et al.
Published: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment
by: Jiang, Zutao, et al.
Published: (2023)
by: Jiang, Zutao, et al.
Published: (2023)
Towards Safe Synthetic Image Generation On the Web: A Multimodal Robust NSFW Defense and Million Scale Dataset
by: Muneer, Muhammad Shahid, et al.
Published: (2025)
by: Muneer, Muhammad Shahid, et al.
Published: (2025)
An Invisible Backdoor Attack Based On Semantic Feature
by: Chen, Yangming
Published: (2024)
by: Chen, Yangming
Published: (2024)
iPad: Iterative Proposal-centric End-to-End Autonomous Driving
by: Guo, Ke, et al.
Published: (2025)
by: Guo, Ke, et al.
Published: (2025)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
Flexiffusion: Training-Free Segment-Wise Neural Architecture Search for Efficient Diffusion Models
by: Huang, Hongtao, et al.
Published: (2025)
by: Huang, Hongtao, et al.
Published: (2025)
Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion Models
by: Chen, Die, et al.
Published: (2025)
by: Chen, Die, et al.
Published: (2025)
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
by: Li, Yuxin, et al.
Published: (2024)
by: Li, Yuxin, et al.
Published: (2024)
BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
by: Li, Feiran, et al.
Published: (2026)
by: Li, Feiran, et al.
Published: (2026)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
by: Gao, Yuansheng, et al.
Published: (2026)
by: Gao, Yuansheng, et al.
Published: (2026)
Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection
by: Zhang, Huixuan, et al.
Published: (2023)
by: Zhang, Huixuan, et al.
Published: (2023)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
by: Zhang, Junzhe, et al.
Published: (2025)
by: Zhang, Junzhe, et al.
Published: (2025)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
by: Poppi, Samuele, et al.
Published: (2023)
by: Poppi, Samuele, et al.
Published: (2023)
Feature-Space Semantic Invariance: Enhanced OOD Detection for Open-Set Domain Generalization
by: Wang, Haoliang, et al.
Published: (2024)
by: Wang, Haoliang, et al.
Published: (2024)
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
by: Liu, Yilian, et al.
Published: (2026)
by: Liu, Yilian, et al.
Published: (2026)
Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models
by: Huang, Zi-Xuan, et al.
Published: (2024)
by: Huang, Zi-Xuan, et al.
Published: (2024)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
by: Kim, Hyunjae, et al.
Published: (2024)
by: Kim, Hyunjae, et al.
Published: (2024)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Audio-Infused Automatic Image Colorization by Exploiting Audio Scene Semantics
by: Zhao, Pengcheng, et al.
Published: (2024)
by: Zhao, Pengcheng, et al.
Published: (2024)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
by: Chen, Haodong, et al.
Published: (2026)
by: Chen, Haodong, et al.
Published: (2026)
Similar Items
-
Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures
by: Liu, Zeyao, et al.
Published: (2026) -
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024) -
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
by: Liu, Xinwei, et al.
Published: (2025) -
CipherDM: Secure Three-Party Inference for Diffusion Model Sampling
by: Zhao, Xin, et al.
Published: (2024) -
Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
by: Zhao, Xin, et al.
Published: (2025)