Bypassing Prompt Guards in Production with Controlled-Release Prompting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fairoze, Jaiden, Garg, Sanjam, Lee, Keewoo, Wang, Mingyuan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Publicly-Detectable Watermarking for Language Models
par: Fairoze, Jaiden, et autres
Publié: (2023)
par: Fairoze, Jaiden, et autres
Publié: (2023)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
par: Fairoze, Jaiden, et autres
Publié: (2025)
par: Fairoze, Jaiden, et autres
Publié: (2025)
SoK: Watermarking for AI-Generated Content
par: Zhao, Xuandong, et autres
Publié: (2024)
par: Zhao, Xuandong, et autres
Publié: (2024)
Black-Box Crypto is Useless for Pseudorandom Codes
par: Garg, Sanjam, et autres
Publié: (2025)
par: Garg, Sanjam, et autres
Publié: (2025)
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
par: Liu, Ting-Chun, et autres
Publié: (2025)
par: Liu, Ting-Chun, et autres
Publié: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
par: Nasr, Milad, et autres
Publié: (2025)
par: Nasr, Milad, et autres
Publié: (2025)
Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing
par: Wahréus, Johan, et autres
Publié: (2025)
par: Wahréus, Johan, et autres
Publié: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
par: Cornacchia, Giandomenico, et autres
Publié: (2024)
par: Cornacchia, Giandomenico, et autres
Publié: (2024)
Are You Using Reliable Graph Prompts? Trojan Prompt Attacks on Graph Neural Networks
par: Lin, Minhua, et autres
Publié: (2024)
par: Lin, Minhua, et autres
Publié: (2024)
Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
par: Nguyen, Quan Minh, et autres
Publié: (2026)
par: Nguyen, Quan Minh, et autres
Publié: (2026)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
par: Zizzo, Giulio, et autres
Publié: (2025)
par: Zizzo, Giulio, et autres
Publié: (2025)
Stealix: Model Stealing via Prompt Evolution
par: Zhuang, Zhixiong, et autres
Publié: (2025)
par: Zhuang, Zhixiong, et autres
Publié: (2025)
Knowledge Return Oriented Prompting (KROP)
par: Martin, Jason, et autres
Publié: (2024)
par: Martin, Jason, et autres
Publié: (2024)
Prompt Obfuscation for Large Language Models
par: Pape, David, et autres
Publié: (2024)
par: Pape, David, et autres
Publié: (2024)
Cape: Context-Aware Prompt Perturbation Mechanism with Differential Privacy
par: Wu, Haoqi, et autres
Publié: (2025)
par: Wu, Haoqi, et autres
Publié: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
par: Yin, Chenlong, et autres
Publié: (2026)
par: Yin, Chenlong, et autres
Publié: (2026)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
par: Zhu, Kaijie, et autres
Publié: (2023)
par: Zhu, Kaijie, et autres
Publié: (2023)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
par: Zou, Wei, et autres
Publié: (2025)
par: Zou, Wei, et autres
Publié: (2025)
Preventing Prompt Injection with Type-Directed Privilege Separation
par: Jacob, Dennis, et autres
Publié: (2025)
par: Jacob, Dennis, et autres
Publié: (2025)
Pr$εε$mpt: Sanitizing Sensitive Prompts for LLMs
par: Chowdhury, Amrita Roy, et autres
Publié: (2025)
par: Chowdhury, Amrita Roy, et autres
Publié: (2025)
Defending Jailbreak Prompts via In-Context Adversarial Game
par: Zhou, Yujun, et autres
Publié: (2024)
par: Zhou, Yujun, et autres
Publié: (2024)
Death by a Thousand Prompts: Open Model Vulnerability Analysis
par: Chang, Amy, et autres
Publié: (2025)
par: Chang, Amy, et autres
Publié: (2025)
Design Patterns for Securing LLM Agents against Prompt Injections
par: Beurer-Kellner, Luca, et autres
Publié: (2025)
par: Beurer-Kellner, Luca, et autres
Publié: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
par: Shi, Chongyang, et autres
Publié: (2025)
par: Shi, Chongyang, et autres
Publié: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
par: Chen, Sizhe, et autres
Publié: (2024)
par: Chen, Sizhe, et autres
Publié: (2024)
Krait: A Backdoor Attack Against Graph Prompt Tuning
par: Song, Ying, et autres
Publié: (2024)
par: Song, Ying, et autres
Publié: (2024)
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
par: Fahey, Ryan
Publié: (2026)
par: Fahey, Ryan
Publié: (2026)
Prompt Stealing Attacks Against Text-to-Image Generation Models
par: Shen, Xinyue, et autres
Publié: (2023)
par: Shen, Xinyue, et autres
Publié: (2023)
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
par: Kang, Mintong, et autres
Publié: (2025)
par: Kang, Mintong, et autres
Publié: (2025)
LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation
par: Lazo, Luis, et autres
Publié: (2026)
par: Lazo, Luis, et autres
Publié: (2026)
Efficient and Adaptable Detection of Malicious LLM Prompts via Bootstrap Aggregation
par: Hassan, Shayan Ali, et autres
Publié: (2026)
par: Hassan, Shayan Ali, et autres
Publié: (2026)
VLMGuard: Defending VLMs against Malicious Prompts via Unlabeled Data
par: Du, Xuefeng, et autres
Publié: (2024)
par: Du, Xuefeng, et autres
Publié: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
par: Xiang, Zhen, et autres
Publié: (2024)
par: Xiang, Zhen, et autres
Publié: (2024)
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
par: Yuan, Shuai, et autres
Publié: (2025)
par: Yuan, Shuai, et autres
Publié: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
par: Hossain, S M Asif, et autres
Publié: (2025)
par: Hossain, S M Asif, et autres
Publié: (2025)
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
par: Pasquini, Dario, et autres
Publié: (2024)
par: Pasquini, Dario, et autres
Publié: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
par: Lu, Yiyang, et autres
Publié: (2026)
par: Lu, Yiyang, et autres
Publié: (2026)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
par: Reddy, Aashray, et autres
Publié: (2025)
par: Reddy, Aashray, et autres
Publié: (2025)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
par: Zhan, Qiusi, et autres
Publié: (2025)
par: Zhan, Qiusi, et autres
Publié: (2025)
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
par: Hsu, Chia-Yi, et autres
Publié: (2026)
par: Hsu, Chia-Yi, et autres
Publié: (2026)
Documents similaires
-
Publicly-Detectable Watermarking for Language Models
par: Fairoze, Jaiden, et autres
Publié: (2023) -
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
par: Fairoze, Jaiden, et autres
Publié: (2025) -
SoK: Watermarking for AI-Generated Content
par: Zhao, Xuandong, et autres
Publié: (2024) -
Black-Box Crypto is Useless for Pseudorandom Codes
par: Garg, Sanjam, et autres
Publié: (2025) -
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
par: Liu, Ting-Chun, et autres
Publié: (2025)