SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Peigui, Tang, Kunsheng, Zhou, Wenbo, Zhang, Weiming, Yu, Nenghai, Zhang, Tianwei, Guo, Qing, Zhang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
Turning Your Strength into Watermark: Watermarking Large Language Model via Knowledge Injection
by: Li, Shuai, et al.
Published: (2023)
by: Li, Shuai, et al.
Published: (2023)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
by: Li, Pengcheng, et al.
Published: (2026)
by: Li, Pengcheng, et al.
Published: (2026)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025)
by: Su, Yanghao, et al.
Published: (2025)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)
by: Su, Yanghao, et al.
Published: (2026)
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
by: Tong, Meng, et al.
Published: (2023)
by: Tong, Meng, et al.
Published: (2023)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRA
by: Feng, Weitao, et al.
Published: (2024)
by: Feng, Weitao, et al.
Published: (2024)
On the Vulnerability of Text Sanitization
by: Tong, Meng, et al.
Published: (2024)
by: Tong, Meng, et al.
Published: (2024)
©Plug-in Authorization for Human Content Copyright Protection in Text-to-Image Model
by: Zhou, Chao, et al.
Published: (2024)
by: Zhou, Chao, et al.
Published: (2024)
Model X-ray:Detecting Backdoored Models via Decision Boundary
by: Su, Yanghao, et al.
Published: (2024)
by: Su, Yanghao, et al.
Published: (2024)
Silent Guardian: Protecting Text from Malicious Exploitation by Large Language Models
by: Zhao, Jiawei, et al.
Published: (2023)
by: Zhao, Jiawei, et al.
Published: (2023)
GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic Security
by: Wang, Xiangkun, et al.
Published: (2025)
by: Wang, Xiangkun, et al.
Published: (2025)
Robust-Wide: Robust Watermarking against Instruction-driven Image Editing
by: Hu, Runyi, et al.
Published: (2024)
by: Hu, Runyi, et al.
Published: (2024)
STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model
by: Qi, Yuang, et al.
Published: (2026)
by: Qi, Yuang, et al.
Published: (2026)
VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts
by: Qi, Peigui, et al.
Published: (2026)
by: Qi, Peigui, et al.
Published: (2026)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Performance-lossless Black-box Model Watermarking
by: Zhao, Na, et al.
Published: (2023)
by: Zhao, Na, et al.
Published: (2023)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
by: Li, Guanlin, et al.
Published: (2024)
by: Li, Guanlin, et al.
Published: (2024)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025)
by: Xin, Yuan, et al.
Published: (2025)
Exploiting Vulnerabilities in Speech Translation Systems through Targeted Adversarial Attacks
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
FoC: Figure out the Cryptographic Functions in Stripped Binaries with LLMs
by: Shang, Xiuwei, et al.
Published: (2024)
by: Shang, Xiuwei, et al.
Published: (2024)
Provably Secure Disambiguating Neural Linguistic Steganography
by: Qi, Yuang, et al.
Published: (2024)
by: Qi, Yuang, et al.
Published: (2024)
A high-capacity linguistic steganography based on entropy-driven rank-token mapping
by: Jiang, Jun, et al.
Published: (2025)
by: Jiang, Jun, et al.
Published: (2025)
EditMark: Watermarking Large Language Models based on Model Editing
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
Provably Secure Public-Key Steganography Based on Admissible Encoding
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
IRCopilot: Automated Incident Response with Large Language Models
by: Lin, Xihuan, et al.
Published: (2025)
by: Lin, Xihuan, et al.
Published: (2025)
Natias: Neuron Attribution based Transferable Image Adversarial Steganography
by: Fan, Zexin, et al.
Published: (2024)
by: Fan, Zexin, et al.
Published: (2024)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Provably Secure Agent Guardrail
by: Wu, Benlong, et al.
Published: (2026)
by: Wu, Benlong, et al.
Published: (2026)
NatGVD: Natural Adversarial Example Attack towards Graph-based Vulnerability Detection
by: Rath, Avilash, et al.
Published: (2025)
by: Rath, Avilash, et al.
Published: (2025)
SiGRRW: A Single-Watermark Robust Reversible Watermarking Framework with Guiding Strategy
by: Xu, Zikai, et al.
Published: (2026)
by: Xu, Zikai, et al.
Published: (2026)
Beyond the Edge of Function: Unraveling the Patterns of Type Recovery in Binary Code
by: Li, Gangyang, et al.
Published: (2025)
by: Li, Gangyang, et al.
Published: (2025)
POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
by: Shao, Yangguang, et al.
Published: (2025)
by: Shao, Yangguang, et al.
Published: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
CamLoPA: A Hidden Wireless Camera Localization Framework via Signal Propagation Path Analysis
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
by: Yang, Zijin, et al.
Published: (2024)
by: Yang, Zijin, et al.
Published: (2024)
VertMark: A Unified Training-Free Robust Watermarking Framework for Vertical Domain Pre-trained Language Models
by: Kong, Cong, et al.
Published: (2026)
by: Kong, Cong, et al.
Published: (2026)
Similar Items
-
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025) -
Turning Your Strength into Watermark: Watermarking Large Language Model via Knowledge Injection
by: Li, Shuai, et al.
Published: (2023) -
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
by: Li, Pengcheng, et al.
Published: (2026) -
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025) -
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)