PromptKeeper: Safeguarding System Prompts for LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jiang, Zhifeng, Jin, Zhihua, He, Guoliang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
par: Shen, Zhili, et autres
Publié: (2024)
par: Shen, Zhili, et autres
Publié: (2024)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
par: Lin, Lixing, et autres
Publié: (2026)
par: Lin, Lixing, et autres
Publié: (2026)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
par: Qi, Xiangyu, et autres
Publié: (2024)
par: Qi, Xiangyu, et autres
Publié: (2024)
Quantized Delta Weight Is Safety Keeper
par: Liu, Yule, et autres
Publié: (2024)
par: Liu, Yule, et autres
Publié: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
par: Wang, Yu, et autres
Publié: (2024)
par: Wang, Yu, et autres
Publié: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
par: Jaiswal, Piyush, et autres
Publié: (2026)
par: Jaiswal, Piyush, et autres
Publié: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
par: Zhang, Zheng, et autres
Publié: (2025)
par: Zhang, Zheng, et autres
Publié: (2025)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
par: Lin, Zheng, et autres
Publié: (2026)
par: Lin, Zheng, et autres
Publié: (2026)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
par: Yang, Yuchen, et autres
Publié: (2024)
par: Yang, Yuchen, et autres
Publié: (2024)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
par: Yeo, Andrew, et autres
Publié: (2025)
par: Yeo, Andrew, et autres
Publié: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
par: Pham, Viet, et autres
Publié: (2025)
par: Pham, Viet, et autres
Publié: (2025)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
par: Levin, Roman, et autres
Publié: (2025)
par: Levin, Roman, et autres
Publié: (2025)
PromptLocate: Localizing Prompt Injection Attacks
par: Jia, Yuqi, et autres
Publié: (2025)
par: Jia, Yuqi, et autres
Publié: (2025)
PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
par: Yu, Jiahao, et autres
Publié: (2024)
par: Yu, Jiahao, et autres
Publié: (2024)
Demo: SGCode: A Flexible Prompt-Optimizing System for Secure Generation of Code
par: Ton, Khiem, et autres
Publié: (2024)
par: Ton, Khiem, et autres
Publié: (2024)
PromptArmor: Simple yet Effective Prompt Injection Defenses
par: Shi, Tianneng, et autres
Publié: (2025)
par: Shi, Tianneng, et autres
Publié: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
par: Wang, Zhilong, et autres
Publié: (2025)
par: Wang, Zhilong, et autres
Publié: (2025)
Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization
par: Verma, Ishaan, et autres
Publié: (2025)
par: Verma, Ishaan, et autres
Publié: (2025)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
par: Wang, Che, et autres
Publié: (2026)
par: Wang, Che, et autres
Publié: (2026)
Assessing Prompt Injection Risks in 200+ Custom GPTs
par: Yu, Jiahao, et autres
Publié: (2023)
par: Yu, Jiahao, et autres
Publié: (2023)
Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
par: Balashov, Andrii, et autres
Publié: (2025)
par: Balashov, Andrii, et autres
Publié: (2025)
Prompt Pirates Need a Map: Stealing Seeds helps Stealing Prompts
par: Mächtle, Felix, et autres
Publié: (2025)
par: Mächtle, Felix, et autres
Publié: (2025)
Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art
par: Trinh, Khoi, et autres
Publié: (2026)
par: Trinh, Khoi, et autres
Publié: (2026)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
par: Liu, Songyang, et autres
Publié: (2026)
par: Liu, Songyang, et autres
Publié: (2026)
Joint Optimization of Prompt Security and System Performance in Edge-Cloud LLM Systems
par: Huang, Haiyang, et autres
Publié: (2025)
par: Huang, Haiyang, et autres
Publié: (2025)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
par: Guo, Xuyang, et autres
Publié: (2025)
par: Guo, Xuyang, et autres
Publié: (2025)
Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
par: Chang, Hongyan, et autres
Publié: (2026)
par: Chang, Hongyan, et autres
Publié: (2026)
Defeating Prompt Injections by Design
par: Debenedetti, Edoardo, et autres
Publié: (2025)
par: Debenedetti, Edoardo, et autres
Publié: (2025)
Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
par: Chang, Zhiyuan, et autres
Publié: (2026)
par: Chang, Zhiyuan, et autres
Publié: (2026)
Safeguarding Large Language Models: A Survey
par: Dong, Yi, et autres
Publié: (2024)
par: Dong, Yi, et autres
Publié: (2024)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
par: Ostermann, Simon, et autres
Publié: (2024)
par: Ostermann, Simon, et autres
Publié: (2024)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
par: Li, Siyuan, et autres
Publié: (2026)
par: Li, Siyuan, et autres
Publié: (2026)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
par: Li, Qinfeng, et autres
Publié: (2025)
par: Li, Qinfeng, et autres
Publié: (2025)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
par: Wang, Zhaoqi, et autres
Publié: (2025)
par: Wang, Zhaoqi, et autres
Publié: (2025)
Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework
par: Zhou, Jiling, et autres
Publié: (2026)
par: Zhou, Jiling, et autres
Publié: (2026)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
par: Cao, Bochuan, et autres
Publié: (2025)
par: Cao, Bochuan, et autres
Publié: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
par: Suo, Xuchen
Publié: (2024)
par: Suo, Xuchen
Publié: (2024)
Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
par: Lee, Donghyun, et autres
Publié: (2024)
par: Lee, Donghyun, et autres
Publié: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
par: Hung, Kuo-Han, et autres
Publié: (2024)
par: Hung, Kuo-Han, et autres
Publié: (2024)
Involuntary Jailbreak: On Self-Prompting Attacks
par: Guo, Yangyang, et autres
Publié: (2025)
par: Guo, Yangyang, et autres
Publié: (2025)
Documents similaires
-
The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
par: Shen, Zhili, et autres
Publié: (2024) -
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
par: Lin, Lixing, et autres
Publié: (2026) -
On Evaluating the Durability of Safeguards for Open-Weight LLMs
par: Qi, Xiangyu, et autres
Publié: (2024) -
Quantized Delta Weight Is Safety Keeper
par: Liu, Yule, et autres
Publié: (2024) -
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
par: Wang, Yu, et autres
Publié: (2024)