An Early Categorization of Prompt Injection Attacks on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rossi, Sippo, Michel, Alisia Marianne, Mukkamala, Raghava Rao, Thatcher, Jason Bennett |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Securing Large Language Models (LLMs) from Prompt Injection Attacks
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023)
von: Yan, Jun, et al.
Veröffentlicht: (2023)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
SPML: A DSL for Defending Language Models Against Prompt Attacks
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
Small Language Models for Curriculum-based Guidance
von: Katharakis, Konstantinos, et al.
Veröffentlicht: (2025)
von: Katharakis, Konstantinos, et al.
Veröffentlicht: (2025)
TFL: Targeted Bit-Flip Attack on Large Language Model
von: Guo, Jingkai, et al.
Veröffentlicht: (2026)
von: Guo, Jingkai, et al.
Veröffentlicht: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025)
von: Formento, Brian, et al.
Veröffentlicht: (2025)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
Prompt Public Large Language Models to Synthesize Data for Private On-device Applications
von: Wu, Shanshan, et al.
Veröffentlicht: (2024)
von: Wu, Shanshan, et al.
Veröffentlicht: (2024)
Prompt Injection Attacks on Large Language Models in Oncology
von: Clusmann, Jan, et al.
Veröffentlicht: (2024)
von: Clusmann, Jan, et al.
Veröffentlicht: (2024)
Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
von: You, Weiqiu, et al.
Veröffentlicht: (2024)
von: You, Weiqiu, et al.
Veröffentlicht: (2024)
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
von: Clop, Cody, et al.
Veröffentlicht: (2024)
von: Clop, Cody, et al.
Veröffentlicht: (2024)
Auditing Prompt Caching in Language Model APIs
von: Gu, Chenchen, et al.
Veröffentlicht: (2025)
von: Gu, Chenchen, et al.
Veröffentlicht: (2025)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
Goal-guided Generative Prompt Injection Attack on Large Language Models
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
PIArena: A Platform for Prompt Injection Evaluation
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
von: An, Bang, et al.
Veröffentlicht: (2024)
von: An, Bang, et al.
Veröffentlicht: (2024)
Model Provenance Testing for Large Language Models
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
Graphene: Infrastructure Security Posture Analysis with AI-generated Attack Graphs
von: Jin, Xin, et al.
Veröffentlicht: (2023)
von: Jin, Xin, et al.
Veröffentlicht: (2023)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Securing Large Language Models (LLMs) from Prompt Injection Attacks
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025) -
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023) -
Defending Against Indirect Prompt Injection Attacks With Spotlighting
von: Hines, Keegan, et al.
Veröffentlicht: (2024) -
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024) -
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)