Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xin, Chen, Xiaojun, Liu, Bingshan, Liu, Zeyao, Zhao, Zhendong, Gu, Xiaoyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures
by: Liu, Zeyao, et al.
Published: (2026)
by: Liu, Zeyao, et al.
Published: (2026)
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
by: Zhao, Xin, et al.
Published: (2025)
by: Zhao, Xin, et al.
Published: (2025)
Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing
by: Chen, Pengzhen, et al.
Published: (2026)
by: Chen, Pengzhen, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
by: Hu, Yuepeng, et al.
Published: (2025)
by: Hu, Yuepeng, et al.
Published: (2025)
CipherDM: Secure Three-Party Inference for Diffusion Model Sampling
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection
by: Liu, Chi, et al.
Published: (2025)
by: Liu, Chi, et al.
Published: (2025)
SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models
by: Liu, Renyang, et al.
Published: (2026)
by: Liu, Renyang, et al.
Published: (2026)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
by: Gao, Sensen, et al.
Published: (2024)
by: Gao, Sensen, et al.
Published: (2024)
Antelope: Potent and Concealed Jailbreak Attack Strategy
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
by: Zhang, Yitong, et al.
Published: (2025)
by: Zhang, Yitong, et al.
Published: (2025)
Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
by: Chin, Zhi-Yi, et al.
Published: (2024)
by: Chin, Zhi-Yi, et al.
Published: (2024)
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples
by: Yang, Yunfei, et al.
Published: (2025)
by: Yang, Yunfei, et al.
Published: (2025)
IPBA: Imperceptible Perturbation Backdoor Attack in Federated Self-Supervised Learning
by: Wang, Jiayao, et al.
Published: (2025)
by: Wang, Jiayao, et al.
Published: (2025)
AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration
by: Liu, Ruhao, et al.
Published: (2026)
by: Liu, Ruhao, et al.
Published: (2026)
ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models
by: Yook, Hyun Jun, et al.
Published: (2025)
by: Yook, Hyun Jun, et al.
Published: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
by: Lyu, Xinqi, et al.
Published: (2025)
by: Lyu, Xinqi, et al.
Published: (2025)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
by: Ba, Zhongjie, et al.
Published: (2023)
by: Ba, Zhongjie, et al.
Published: (2023)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
by: Jia, Xiaojun, et al.
Published: (2025)
by: Jia, Xiaojun, et al.
Published: (2025)
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
by: Liu, Mingrui, et al.
Published: (2025)
by: Liu, Mingrui, et al.
Published: (2025)
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
by: Wu, Xiaodong, et al.
Published: (2025)
by: Wu, Xiaodong, et al.
Published: (2025)
Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning
by: Komaty, Alain, et al.
Published: (2025)
by: Komaty, Alain, et al.
Published: (2025)
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
VGS-ATD: Robust Distributed Learning for Multi-Label Medical Image Classification Under Heterogeneous and Imbalanced Conditions
by: Zhao, Zehui, et al.
Published: (2025)
by: Zhao, Zehui, et al.
Published: (2025)
ZeroPur: Succinct Training-Free Adversarial Purification
by: Liu, Erhu, et al.
Published: (2024)
by: Liu, Erhu, et al.
Published: (2024)
Anti-Tamper Protection for Unauthorized Individual Image Generation
by: Li, Zelin, et al.
Published: (2025)
by: Li, Zelin, et al.
Published: (2025)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security
by: Fan, Yihe, et al.
Published: (2024)
by: Fan, Yihe, et al.
Published: (2024)
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
by: Ren, Huali, et al.
Published: (2024)
by: Ren, Huali, et al.
Published: (2024)
LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors
by: Lu, Jiajie, et al.
Published: (2025)
by: Lu, Jiajie, et al.
Published: (2025)
Robust Anti-Backdoor Instruction Tuning in LVLMs
by: Xun, Yuan, et al.
Published: (2025)
by: Xun, Yuan, et al.
Published: (2025)
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
by: Liu, Yule, et al.
Published: (2026)
by: Liu, Yule, et al.
Published: (2026)
Is Diffusion Model Safe? Severe Data Leakage via Gradient-Guided Diffusion Model
by: Meng, Jiayang, et al.
Published: (2024)
by: Meng, Jiayang, et al.
Published: (2024)
Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling
by: Chen, Pengzhen, et al.
Published: (2026)
by: Chen, Pengzhen, et al.
Published: (2026)
Similar Items
-
Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures
by: Liu, Zeyao, et al.
Published: (2026) -
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
by: Zhao, Xin, et al.
Published: (2025) -
Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing
by: Chen, Pengzhen, et al.
Published: (2026) -
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025) -
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
by: Hu, Yuepeng, et al.
Published: (2025)