PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Yuan, Lingzhi, Li, Xinfeng, Xu, Chejian, Tao, Guanhong, Jia, Xiaojun, Huang, Yihao, Dong, Wei, Liu, Yang, Li, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
por: Wu, Yixin, et al.
Publicado: (2023)
por: Wu, Yixin, et al.
Publicado: (2023)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
por: Zhao, Shiqian, et al.
Publicado: (2025)
por: Zhao, Shiqian, et al.
Publicado: (2025)
X-Guard: Multilingual Guard Agent for Content Moderation
por: Upadhayay, Bibek, et al.
Publicado: (2025)
por: Upadhayay, Bibek, et al.
Publicado: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
por: Kang, Mintong, et al.
Publicado: (2025)
por: Kang, Mintong, et al.
Publicado: (2025)
Bypassing Prompt Guards in Production with Controlled-Release Prompting
por: Fairoze, Jaiden, et al.
Publicado: (2025)
por: Fairoze, Jaiden, et al.
Publicado: (2025)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
por: Chen, Yulin, et al.
Publicado: (2026)
por: Chen, Yulin, et al.
Publicado: (2026)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
por: Lyu, Xinqi, et al.
Publicado: (2025)
por: Lyu, Xinqi, et al.
Publicado: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
por: Teng, Ma, et al.
Publicado: (2024)
por: Teng, Ma, et al.
Publicado: (2024)
LLMGuard: Guarding Against Unsafe LLM Behavior
por: Goyal, Shubh, et al.
Publicado: (2024)
por: Goyal, Shubh, et al.
Publicado: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
por: Liu, Hongfu, et al.
Publicado: (2024)
por: Liu, Hongfu, et al.
Publicado: (2024)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
por: He, Yu, et al.
Publicado: (2026)
por: He, Yu, et al.
Publicado: (2026)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026)
por: Wang, Peiran, et al.
Publicado: (2026)
Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies
por: Wang, Peiran, et al.
Publicado: (2024)
por: Wang, Peiran, et al.
Publicado: (2024)
MacPrompt: Maraconic-guided Jailbreak against Text-to-Image Models
por: Ye, Xi, et al.
Publicado: (2026)
por: Ye, Xi, et al.
Publicado: (2026)
Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
por: Nguyen, Quan Minh, et al.
Publicado: (2026)
por: Nguyen, Quan Minh, et al.
Publicado: (2026)
CourtGuard: A Local, Multiagent Prompt Injection Classifier
por: Wu, Isaac, et al.
Publicado: (2025)
por: Wu, Isaac, et al.
Publicado: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
por: Xiang, Yuxiao, et al.
Publicado: (2025)
por: Xiang, Yuxiao, et al.
Publicado: (2025)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
por: Li, Jun, et al.
Publicado: (2026)
por: Li, Jun, et al.
Publicado: (2026)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
por: Kang, Mintong, et al.
Publicado: (2024)
por: Kang, Mintong, et al.
Publicado: (2024)
RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation
por: Tafreshian, Benyamin
Publicado: (2025)
por: Tafreshian, Benyamin
Publicado: (2025)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
por: Gong, Guangyu, et al.
Publicado: (2026)
por: Gong, Guangyu, et al.
Publicado: (2026)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
por: Langiu, Alessio
Publicado: (2026)
por: Langiu, Alessio
Publicado: (2026)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
por: Lin, Lixing, et al.
Publicado: (2026)
por: Lin, Lixing, et al.
Publicado: (2026)
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
por: Yang, Zhuochen, et al.
Publicado: (2025)
por: Yang, Zhuochen, et al.
Publicado: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
por: Meng, Xiangtao, et al.
Publicado: (2025)
por: Meng, Xiangtao, et al.
Publicado: (2025)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
por: Zhao, Wei, et al.
Publicado: (2026)
por: Zhao, Wei, et al.
Publicado: (2026)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
por: Du, Mengyao, et al.
Publicado: (2026)
por: Du, Mengyao, et al.
Publicado: (2026)
Prompt Stealing Attacks Against Text-to-Image Generation Models
por: Shen, Xinyue, et al.
Publicado: (2023)
por: Shen, Xinyue, et al.
Publicado: (2023)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
por: Wang, Yiming, et al.
Publicado: (2024)
por: Wang, Yiming, et al.
Publicado: (2024)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
por: Gao, Sensen, et al.
Publicado: (2024)
por: Gao, Sensen, et al.
Publicado: (2024)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
por: Huang, Xun, et al.
Publicado: (2026)
por: Huang, Xun, et al.
Publicado: (2026)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
por: Zhu, Zhenhao, et al.
Publicado: (2026)
por: Zhu, Zhenhao, et al.
Publicado: (2026)
SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
por: Yang, Wenyuan, et al.
Publicado: (2025)
por: Yang, Wenyuan, et al.
Publicado: (2025)
ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
por: Zhuang, Zhixiong, et al.
Publicado: (2025)
por: Zhuang, Zhixiong, et al.
Publicado: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
Alleviating the Fear of Losing Alignment in LLM Fine-tuning
por: Yang, Kang, et al.
Publicado: (2025)
por: Yang, Kang, et al.
Publicado: (2025)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
por: Schwartz, Daniel, et al.
Publicado: (2025)
por: Schwartz, Daniel, et al.
Publicado: (2025)
Ejemplares similares
-
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
por: Zhang, Xiaoyu, et al.
Publicado: (2023) -
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
por: Wu, Yixin, et al.
Publicado: (2023) -
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
por: Zhao, Shiqian, et al.
Publicado: (2025) -
X-Guard: Multilingual Guard Agent for Content Moderation
por: Upadhayay, Bibek, et al.
Publicado: (2025) -
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
por: Kang, Mintong, et al.
Publicado: (2025)