Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
Fuente:
arXiv
Guardado en:
| Autores principales: | Meng, Xiangtao, Dong, Yingkai, Yu, Ning, Wang, Li, Li, Zheng, Guo, Shanqing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment
por: Wang, Li, et al.
Publicado: (2025)
por: Wang, Li, et al.
Publicado: (2025)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
por: Dong, Yingkai, et al.
Publicado: (2024)
por: Dong, Yingkai, et al.
Publicado: (2024)
AVA: Inconspicuous Attribute Variation-based Adversarial Attack bypassing DeepFake Detection
por: Meng, Xiangtao, et al.
Publicado: (2023)
por: Meng, Xiangtao, et al.
Publicado: (2023)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
por: Wang, Li, et al.
Publicado: (2026)
por: Wang, Li, et al.
Publicado: (2026)
DCMI: A Differential Calibration Membership Inference Attack Against Retrieval-Augmented Generation
por: Gao, Xinyu, et al.
Publicado: (2025)
por: Gao, Xinyu, et al.
Publicado: (2025)
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
por: Qu, Yiting, et al.
Publicado: (2024)
por: Qu, Yiting, et al.
Publicado: (2024)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
por: Xiang, Yuxiao, et al.
Publicado: (2025)
por: Xiang, Yuxiao, et al.
Publicado: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
por: Yuan, Lingzhi, et al.
Publicado: (2025)
por: Yuan, Lingzhi, et al.
Publicado: (2025)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
por: Ba, Zhongjie, et al.
Publicado: (2023)
por: Ba, Zhongjie, et al.
Publicado: (2023)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
por: Li, Jun, et al.
Publicado: (2026)
por: Li, Jun, et al.
Publicado: (2026)
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
por: Zhang, Yitong, et al.
Publicado: (2025)
por: Zhang, Yitong, et al.
Publicado: (2025)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
por: Wu, Yixin, et al.
Publicado: (2023)
por: Wu, Yixin, et al.
Publicado: (2023)
FaceSwapGuard: Safeguarding Facial Privacy from DeepFake Threats through Identity Obfuscation
por: Wang, Li, et al.
Publicado: (2025)
por: Wang, Li, et al.
Publicado: (2025)
Membership Inference Attack Against Masked Image Modeling
por: Li, Zheng, et al.
Publicado: (2024)
por: Li, Zheng, et al.
Publicado: (2024)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
por: Zhang, Chenyu, et al.
Publicado: (2025)
por: Zhang, Chenyu, et al.
Publicado: (2025)
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
por: Qi, Peigui, et al.
Publicado: (2025)
por: Qi, Peigui, et al.
Publicado: (2025)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
por: Liu, Yule, et al.
Publicado: (2026)
por: Liu, Yule, et al.
Publicado: (2026)
From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
por: Meng, Xiangtao, et al.
Publicado: (2025)
por: Meng, Xiangtao, et al.
Publicado: (2025)
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy
por: Zhai, Shengfang, et al.
Publicado: (2024)
por: Zhai, Shengfang, et al.
Publicado: (2024)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
por: Guo, Yangyang, et al.
Publicado: (2024)
por: Guo, Yangyang, et al.
Publicado: (2024)
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
por: Leng, Ye, et al.
Publicado: (2026)
por: Leng, Ye, et al.
Publicado: (2026)
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
por: Xiong, Yuan, et al.
Publicado: (2025)
por: Xiong, Yuan, et al.
Publicado: (2025)
The Structural Safety Generalization Problem
por: Broomfield, Julius, et al.
Publicado: (2025)
por: Broomfield, Julius, et al.
Publicado: (2025)
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
por: Li, Boheng, et al.
Publicado: (2025)
por: Li, Boheng, et al.
Publicado: (2025)
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
por: Guo, Ji, et al.
Publicado: (2024)
por: Guo, Ji, et al.
Publicado: (2024)
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
por: Miao, Yibo, et al.
Publicado: (2024)
por: Miao, Yibo, et al.
Publicado: (2024)
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
por: Dang, Pucheng, et al.
Publicado: (2024)
por: Dang, Pucheng, et al.
Publicado: (2024)
ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation
por: Teng, Qianrui, et al.
Publicado: (2025)
por: Teng, Qianrui, et al.
Publicado: (2025)
Towards Understanding Unsafe Video Generation
por: Pang, Yan, et al.
Publicado: (2024)
por: Pang, Yan, et al.
Publicado: (2024)
Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models
por: Meng, Xiangtao, et al.
Publicado: (2026)
por: Meng, Xiangtao, et al.
Publicado: (2026)
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
por: Zhao, Shiqian, et al.
Publicado: (2025)
por: Zhao, Shiqian, et al.
Publicado: (2025)
Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling
por: Li, Zida, et al.
Publicado: (2026)
por: Li, Zida, et al.
Publicado: (2026)
Responsible Diffusion: A Comprehensive Survey on Safety, Ethics, and Trust in Diffusion Models
por: Wei, Kang, et al.
Publicado: (2025)
por: Wei, Kang, et al.
Publicado: (2025)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
por: Luo, Zeren, et al.
Publicado: (2025)
por: Luo, Zeren, et al.
Publicado: (2025)
Unveiling and Mitigating Memorization in Text-to-image Diffusion Models through Cross Attention
por: Ren, Jie, et al.
Publicado: (2024)
por: Ren, Jie, et al.
Publicado: (2024)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
por: Ye, Mang, et al.
Publicado: (2025)
por: Ye, Mang, et al.
Publicado: (2025)
SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
por: Rong, Xuankun, et al.
Publicado: (2025)
por: Rong, Xuankun, et al.
Publicado: (2025)
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
por: Li, Xinfeng, et al.
Publicado: (2024)
por: Li, Xinfeng, et al.
Publicado: (2024)
Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models
por: Ahn, Jaesin, et al.
Publicado: (2025)
por: Ahn, Jaesin, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment
por: Wang, Li, et al.
Publicado: (2025) -
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
por: Dong, Yingkai, et al.
Publicado: (2024) -
AVA: Inconspicuous Attribute Variation-based Adversarial Attack bypassing DeepFake Detection
por: Meng, Xiangtao, et al.
Publicado: (2023) -
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
por: Wang, Li, et al.
Publicado: (2026) -
DCMI: A Differential Calibration Membership Inference Attack Against Retrieval-Augmented Generation
por: Gao, Xinyu, et al.
Publicado: (2025)