SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hongxiang, He, Yifeng, Chen, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
by: Liu, Mingrui, et al.
Published: (2025)
by: Liu, Mingrui, et al.
Published: (2025)
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
by: Zhang, Chenyu, et al.
Published: (2024)
by: Zhang, Chenyu, et al.
Published: (2024)
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
by: Li, Xinfeng, et al.
Published: (2024)
by: Li, Xinfeng, et al.
Published: (2024)
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
by: Dang, Pucheng, et al.
Published: (2024)
by: Dang, Pucheng, et al.
Published: (2024)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
by: Zhang, Zhuomeng, et al.
Published: (2024)
by: Zhang, Zhuomeng, et al.
Published: (2024)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
by: Qi, Peigui, et al.
Published: (2025)
by: Qi, Peigui, et al.
Published: (2025)
Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
Metaphor-based Jailbreak Attacks on Text-to-Image Models
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
by: Xu, Peiyang, et al.
Published: (2025)
by: Xu, Peiyang, et al.
Published: (2025)
Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
by: Li, Leyang, et al.
Published: (2025)
by: Li, Leyang, et al.
Published: (2025)
T2UE: Generating Unlearnable Examples from Text Descriptions
by: Ma, Xingjun, et al.
Published: (2025)
by: Ma, Xingjun, et al.
Published: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
CipherDM: Secure Three-Party Inference for Diffusion Model Sampling
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
by: Lyu, Xinqi, et al.
Published: (2025)
by: Lyu, Xinqi, et al.
Published: (2025)
Security Risk of Misalignment between Text and Image in Multi-modal Model
by: Wang, Xiaosen, et al.
Published: (2025)
by: Wang, Xiaosen, et al.
Published: (2025)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
by: Lee, Junhoo, et al.
Published: (2026)
by: Lee, Junhoo, et al.
Published: (2026)
SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models
by: Liu, Renyang, et al.
Published: (2026)
by: Liu, Renyang, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
ICAS: Detecting Training Data from Autoregressive Image Generative Models
by: Yu, Hongyao, et al.
Published: (2025)
by: Yu, Hongyao, et al.
Published: (2025)
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
by: Pu, Jiayue, et al.
Published: (2026)
by: Pu, Jiayue, et al.
Published: (2026)
MetaCloak: Preventing Unauthorized Subject-driven Text-to-image Diffusion-based Synthesis via Meta-learning
by: Liu, Yixin, et al.
Published: (2023)
by: Liu, Yixin, et al.
Published: (2023)
Gradient Inversion of Federated Diffusion Models
by: Huang, Jiyue, et al.
Published: (2024)
by: Huang, Jiyue, et al.
Published: (2024)
Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Backdoor Attacks against Image-to-Image Networks
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Differentially Private Fine-Tuning of Diffusion Models
by: Tsai, Yu-Lin, et al.
Published: (2024)
by: Tsai, Yu-Lin, et al.
Published: (2024)
Text is All You Need for Vision-Language Model Jailbreaking
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models
by: Alam, Mohammed Talha, et al.
Published: (2025)
by: Alam, Mohammed Talha, et al.
Published: (2025)
CopyrightMeter: Revisiting Copyright Protection in Text-to-image Models
by: Xu, Naen, et al.
Published: (2024)
by: Xu, Naen, et al.
Published: (2024)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
by: Xu, Shuhan, et al.
Published: (2026)
by: Xu, Shuhan, et al.
Published: (2026)
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models
by: Li, Guo, et al.
Published: (2025)
by: Li, Guo, et al.
Published: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
©Plug-in Authorization for Human Content Copyright Protection in Text-to-Image Model
by: Zhou, Chao, et al.
Published: (2024)
by: Zhou, Chao, et al.
Published: (2024)
Similar Items
-
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
by: Liu, Mingrui, et al.
Published: (2025) -
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
by: Zhang, Chenyu, et al.
Published: (2024) -
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
by: Li, Xinfeng, et al.
Published: (2024) -
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
by: Dang, Pucheng, et al.
Published: (2024) -
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
by: Zhang, Zhuomeng, et al.
Published: (2024)