ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhuang, Zhixiong, Nicolae, Maria-Irina, Wang, Hui-Po, Fritz, Mario |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Stealix: Model Stealing via Prompt Evolution
por: Zhuang, Zhixiong, et al.
Publicado: (2025)
por: Zhuang, Zhixiong, et al.
Publicado: (2025)
Stealthy Imitation: Reward-guided Environment-free Policy Stealing
por: Zhuang, Zhixiong, et al.
Publicado: (2024)
por: Zhuang, Zhixiong, et al.
Publicado: (2024)
System Prompt Extraction Attacks and Defenses in Large Language Models
por: Das, Badhan Chandra, et al.
Publicado: (2025)
por: Das, Badhan Chandra, et al.
Publicado: (2025)
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
por: Shen, Yaling, et al.
Publicado: (2025)
por: Shen, Yaling, et al.
Publicado: (2025)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
por: Yang, Yong, et al.
Publicado: (2024)
por: Yang, Yong, et al.
Publicado: (2024)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
por: Cui, Yu, et al.
Publicado: (2025)
por: Cui, Yu, et al.
Publicado: (2025)
PINA: Prompt Injection Attack against Navigation Agents
por: Liu, Jiani, et al.
Publicado: (2026)
por: Liu, Jiani, et al.
Publicado: (2026)
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
por: Qu, Wenjie, et al.
Publicado: (2025)
por: Qu, Wenjie, et al.
Publicado: (2025)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
por: Zhao, Shiqian, et al.
Publicado: (2025)
por: Zhao, Shiqian, et al.
Publicado: (2025)
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
por: Peng, Benji, et al.
Publicado: (2024)
por: Peng, Benji, et al.
Publicado: (2024)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
por: Wu, Fangzhou, et al.
Publicado: (2024)
por: Wu, Fangzhou, et al.
Publicado: (2024)
Securing AI Agents Against Prompt Injection Attacks
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
Design Patterns for Securing LLM Agents against Prompt Injections
por: Beurer-Kellner, Luca, et al.
Publicado: (2025)
por: Beurer-Kellner, Luca, et al.
Publicado: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
por: Yin, Yu, et al.
Publicado: (2026)
por: Yin, Yu, et al.
Publicado: (2026)
A Critical Evaluation of Defenses against Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
por: Ma, Jiachen, et al.
Publicado: (2024)
por: Ma, Jiachen, et al.
Publicado: (2024)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
por: Hui, Bo, et al.
Publicado: (2024)
por: Hui, Bo, et al.
Publicado: (2024)
Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
por: Nguyen, Quan Minh, et al.
Publicado: (2026)
por: Nguyen, Quan Minh, et al.
Publicado: (2026)
Strengthening Polymorphic Prompt Assembling: Dynamic Separator Generation Against Emerging Prompt Injection Attacks
por: Dorzhiev, Nima, et al.
Publicado: (2026)
por: Dorzhiev, Nima, et al.
Publicado: (2026)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
Are You Using Reliable Graph Prompts? Trojan Prompt Attacks on Graph Neural Networks
por: Lin, Minhua, et al.
Publicado: (2024)
por: Lin, Minhua, et al.
Publicado: (2024)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2026)
por: Jia, Yuqi, et al.
Publicado: (2026)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
por: Wang, Jiawen, et al.
Publicado: (2025)
por: Wang, Jiawen, et al.
Publicado: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
por: Evtimov, Ivan, et al.
Publicado: (2025)
por: Evtimov, Ivan, et al.
Publicado: (2025)
PromptKeeper: Safeguarding System Prompts for LLMs
por: Jiang, Zhifeng, et al.
Publicado: (2024)
por: Jiang, Zhifeng, et al.
Publicado: (2024)
Prompt Injection Attack to Tool Selection in LLM Agents
por: Shi, Jiawen, et al.
Publicado: (2025)
por: Shi, Jiawen, et al.
Publicado: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
MacPrompt: Maraconic-guided Jailbreak against Text-to-Image Models
por: Ye, Xi, et al.
Publicado: (2026)
por: Ye, Xi, et al.
Publicado: (2026)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
por: Wang, Mengxiao, et al.
Publicado: (2025)
por: Wang, Mengxiao, et al.
Publicado: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
por: Li, Yucheng, et al.
Publicado: (2025)
por: Li, Yucheng, et al.
Publicado: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
por: Guo, Yangyang, et al.
Publicado: (2025)
por: Guo, Yangyang, et al.
Publicado: (2025)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
por: Li, Yanjie, et al.
Publicado: (2025)
por: Li, Yanjie, et al.
Publicado: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025)
por: Zizzo, Giulio, et al.
Publicado: (2025)
Ejemplares similares
-
Stealix: Model Stealing via Prompt Evolution
por: Zhuang, Zhixiong, et al.
Publicado: (2025) -
Stealthy Imitation: Reward-guided Environment-free Policy Stealing
por: Zhuang, Zhixiong, et al.
Publicado: (2024) -
System Prompt Extraction Attacks and Defenses in Large Language Models
por: Das, Badhan Chandra, et al.
Publicado: (2025) -
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
por: Shen, Yaling, et al.
Publicado: (2025) -
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
por: Yang, Yong, et al.
Publicado: (2024)