Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Jiwoo, Na, Byeonghu, Kang, Mina, Choi, Wonhyeok, Moon, Il-Chul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
by: Na, Byeonghu, et al.
Published: (2025)
by: Na, Byeonghu, et al.
Published: (2025)
Dirichlet-based Per-Sample Weighting by Transition Matrix for Noisy Label Learning
by: Bae, HeeSun, et al.
Published: (2024)
by: Bae, HeeSun, et al.
Published: (2024)
Semantic Guidance Tuning for Text-To-Image Diffusion Models
by: Kang, Hyun, et al.
Published: (2023)
by: Kang, Hyun, et al.
Published: (2023)
Training Unbiased Diffusion Models From Biased Dataset
by: Kim, Yeongmin, et al.
Published: (2024)
by: Kim, Yeongmin, et al.
Published: (2024)
Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
by: Na, Byeonghu, et al.
Published: (2025)
by: Na, Byeonghu, et al.
Published: (2025)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Distillation of Large Language Models via Concrete Score Matching
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models
by: Koma, Arian Komaei, et al.
Published: (2026)
by: Koma, Arian Komaei, et al.
Published: (2026)
Depth-discriminative Metric Learning for Monocular 3D Object Detection
by: Choi, Wonhyeok, et al.
Published: (2024)
by: Choi, Wonhyeok, et al.
Published: (2024)
Distilling Dataset into Neural Field
by: Shin, Donghyeok, et al.
Published: (2025)
by: Shin, Donghyeok, et al.
Published: (2025)
Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
by: Kim, Yeongmin, et al.
Published: (2026)
by: Kim, Yeongmin, et al.
Published: (2026)
Personalized Safety Alignment for Text-to-Image Diffusion Models
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Minority-Focused Text-to-Image Generation via Prompt Optimization
by: Um, Soobin, et al.
Published: (2024)
by: Um, Soobin, et al.
Published: (2024)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
by: Peng, Xingkai, et al.
Published: (2025)
by: Peng, Xingkai, et al.
Published: (2025)
Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning
by: Moon, Saemi, et al.
Published: (2024)
by: Moon, Saemi, et al.
Published: (2024)
Distribution-Level Feature Distancing for Machine Unlearning: Towards a Better Trade-off Between Model Utility and Forgetting
by: Choi, Dasol, et al.
Published: (2024)
by: Choi, Dasol, et al.
Published: (2024)
CharDiff-LP: A Diffusion Model with Character-Level Guidance for License Plate Image Restoration
by: Na, Kihyun, et al.
Published: (2025)
by: Na, Kihyun, et al.
Published: (2025)
Don't Play Favorites: Minority Guidance for Diffusion Models
by: Um, Soobin, et al.
Published: (2023)
by: Um, Soobin, et al.
Published: (2023)
Counting Guidance for High Fidelity Text-to-Image Synthesis
by: Kang, Wonjun, et al.
Published: (2023)
by: Kang, Wonjun, et al.
Published: (2023)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
DiffGuard: Text-Based Safety Checker for Diffusion Models
by: Khader, Massine El, et al.
Published: (2024)
by: Khader, Massine El, et al.
Published: (2024)
PromptLoop: Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
by: Lee, Suhyeon, et al.
Published: (2025)
by: Lee, Suhyeon, et al.
Published: (2025)
WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single Image
by: Park, Jiwoo, et al.
Published: (2025)
by: Park, Jiwoo, et al.
Published: (2025)
Steering Guidance for Personalized Text-to-Image Diffusion Models
by: Park, Sunghyun, et al.
Published: (2025)
by: Park, Sunghyun, et al.
Published: (2025)
Latent Schrodinger Bridge: Prompting Latent Diffusion for Fast Unpaired Image-to-Image Translation
by: Kim, Jeongsol, et al.
Published: (2024)
by: Kim, Jeongsol, et al.
Published: (2024)
Aligning Text to Image in Diffusion Models is Easier Than You Think
by: Lee, Jaa-Yeon, et al.
Published: (2025)
by: Lee, Jaa-Yeon, et al.
Published: (2025)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
by: Lee, Jaa-Yeon, et al.
Published: (2026)
by: Lee, Jaa-Yeon, et al.
Published: (2026)
Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models
by: Moon, Saemi, et al.
Published: (2026)
by: Moon, Saemi, et al.
Published: (2026)
Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
by: Zou, Xiaotian
Published: (2025)
by: Zou, Xiaotian
Published: (2025)
Generative Unlearning for Any Identity
by: Seo, Juwon, et al.
Published: (2024)
by: Seo, Juwon, et al.
Published: (2024)
CFG++: Manifold-constrained Classifier Free Guidance for Diffusion Models
by: Chung, Hyungjin, et al.
Published: (2024)
by: Chung, Hyungjin, et al.
Published: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Dynamic VLM-Guided Negative Prompting for Diffusion Models
by: Chang, Hoyeon, et al.
Published: (2025)
by: Chang, Hoyeon, et al.
Published: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
by: Schneider, Moritz, et al.
Published: (2024)
by: Schneider, Moritz, et al.
Published: (2024)
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
by: Fan, Zezhong, et al.
Published: (2024)
by: Fan, Zezhong, et al.
Published: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Similar Items
-
Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
by: Na, Byeonghu, et al.
Published: (2025) -
Dirichlet-based Per-Sample Weighting by Transition Matrix for Noisy Label Learning
by: Bae, HeeSun, et al.
Published: (2024) -
Semantic Guidance Tuning for Text-To-Image Diffusion Models
by: Kang, Hyun, et al.
Published: (2023) -
Training Unbiased Diffusion Models From Biased Dataset
by: Kim, Yeongmin, et al.
Published: (2024) -
Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
by: Na, Byeonghu, et al.
Published: (2025)