SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinfeng, Yang, Yuchen, Deng, Jiangyi, Yan, Chen, Chen, Yanjiao, Ji, Xiaoyu, Xu, Wenyuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
by: Li, Xinfeng, et al.
Published: (2025)
by: Li, Xinfeng, et al.
Published: (2025)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
by: Qiu, Huming, et al.
Published: (2024)
by: Qiu, Huming, et al.
Published: (2024)
RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer
by: Deng, Jiangyi, et al.
Published: (2024)
by: Deng, Jiangyi, et al.
Published: (2024)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
by: Wang, Zilong, et al.
Published: (2025)
by: Wang, Zilong, et al.
Published: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)
by: Zhong, Yinan, et al.
Published: (2025)
SafeEar: Content Privacy-Preserving Audio Deepfake Detection
by: Li, Xinfeng, et al.
Published: (2024)
by: Li, Xinfeng, et al.
Published: (2024)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
by: Wu, Jialin, et al.
Published: (2024)
by: Wu, Jialin, et al.
Published: (2024)
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
by: Zhang, Junbo, et al.
Published: (2025)
by: Zhang, Junbo, et al.
Published: (2025)
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
by: Meng, Wenlong, et al.
Published: (2025)
by: Meng, Wenlong, et al.
Published: (2025)
A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts
by: Chen, Yingquan, et al.
Published: (2025)
by: Chen, Yingquan, et al.
Published: (2025)
A General Pseudonymization Framework for Cloud-Based LLMs: Replacing Privacy Information in Controlled Text Generation
by: Hou, Shilong, et al.
Published: (2025)
by: Hou, Shilong, et al.
Published: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
by: Singh, Himanshu, et al.
Published: (2026)
by: Singh, Himanshu, et al.
Published: (2026)
Secret-Protected Evolution for Differentially Private Synthetic Text Generation
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
From Retrieval to Reasoning: A Framework for Cyber Threat Intelligence NER with Explicit and Adaptive Instructions
by: Peng, Jiaren, et al.
Published: (2025)
by: Peng, Jiaren, et al.
Published: (2025)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
by: Cui, Shiyao, et al.
Published: (2023)
by: Cui, Shiyao, et al.
Published: (2023)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
Watermarking Large Language Models and the Generated Content: Opportunities and Challenges
by: Zhang, Ruisi, et al.
Published: (2024)
by: Zhang, Ruisi, et al.
Published: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
by: Dai, Fangqi, et al.
Published: (2025)
by: Dai, Fangqi, et al.
Published: (2025)
A Content-Preserving Secure Linguistic Steganography
by: Xiang, Lingyun, et al.
Published: (2025)
by: Xiang, Lingyun, et al.
Published: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems
by: Bodea, Andreea-Elena, et al.
Published: (2026)
by: Bodea, Andreea-Elena, et al.
Published: (2026)
Adversarial Text Generation with Dynamic Contextual Perturbation
by: Waghela, Hetvi, et al.
Published: (2025)
by: Waghela, Hetvi, et al.
Published: (2025)
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
by: Zhou, Ying, et al.
Published: (2024)
by: Zhou, Ying, et al.
Published: (2024)
EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
by: Meisenbacher, Stephen, et al.
Published: (2025)
by: Meisenbacher, Stephen, et al.
Published: (2025)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
by: Wang, Lionel Z., et al.
Published: (2026)
by: Wang, Lionel Z., et al.
Published: (2026)
False Reality: Uncovering Sensor-induced Human-VR Interaction Vulnerability
by: Jiang, Yancheng, et al.
Published: (2025)
by: Jiang, Yancheng, et al.
Published: (2025)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
by: Guo, Ruoqi, et al.
Published: (2026)
by: Guo, Ruoqi, et al.
Published: (2026)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
by: Li, Yuanfan, et al.
Published: (2026)
by: Li, Yuanfan, et al.
Published: (2026)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models
by: Ahn, Jaesin, et al.
Published: (2025)
by: Ahn, Jaesin, et al.
Published: (2025)
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
by: Zhang, Jiankun, et al.
Published: (2025)
by: Zhang, Jiankun, et al.
Published: (2025)
TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity
by: Cao, Xi, et al.
Published: (2024)
by: Cao, Xi, et al.
Published: (2024)
Similar Items
-
Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
by: Li, Xinfeng, et al.
Published: (2025) -
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
by: Qiu, Huming, et al.
Published: (2024) -
RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer
by: Deng, Jiangyi, et al.
Published: (2024) -
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
by: Wang, Zilong, et al.
Published: (2025) -
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)