Prompt Injection Evaluations: Refusal Boundary Instability and Artifact-Dependent Compliance in GPT-4-Series Models
Fuente:
arXiv
Salvato in:
| Autore principale: | Heverin, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Breaking the Prompt Wall (I): A Real-World Case Study of Attacking ChatGPT via Lightweight Prompt Injection
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Privacy Artifact ConnecTor (PACT): Embedding Enterprise Artifacts for Compliance AI Agents
di: Fang, Chenhao, et al.
Pubblicazione: (2025)
di: Fang, Chenhao, et al.
Pubblicazione: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
di: Deep, Priyal, et al.
Pubblicazione: (2026)
di: Deep, Priyal, et al.
Pubblicazione: (2026)
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals
di: Akinrele, Akindoyin, et al.
Pubblicazione: (2026)
di: Akinrele, Akindoyin, et al.
Pubblicazione: (2026)
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
di: Wu, Tongxi, et al.
Pubblicazione: (2026)
di: Wu, Tongxi, et al.
Pubblicazione: (2026)
PromptShield: Deployable Detection for Prompt Injection Attacks
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
A Novel Evaluation Framework for Assessing Resilience Against Prompt Injection Attacks in Large Language Models
di: Yip, Daniel Wankit, et al.
Pubblicazione: (2024)
di: Yip, Daniel Wankit, et al.
Pubblicazione: (2024)
Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models
di: Ganiuly, Daniyal, et al.
Pubblicazione: (2025)
di: Ganiuly, Daniyal, et al.
Pubblicazione: (2025)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
di: Wang, Mengxiao, et al.
Pubblicazione: (2025)
di: Wang, Mengxiao, et al.
Pubblicazione: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
di: Yin, Yu, et al.
Pubblicazione: (2026)
di: Yin, Yu, et al.
Pubblicazione: (2026)
Defending Against Prompt Injection with DataFilter
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models
di: Sandoval, Gustavo, et al.
Pubblicazione: (2025)
di: Sandoval, Gustavo, et al.
Pubblicazione: (2025)
Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding
di: Pathade, Chetan
Pubblicazione: (2025)
di: Pathade, Chetan
Pubblicazione: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
di: Wang, Reachal, et al.
Pubblicazione: (2025)
di: Wang, Reachal, et al.
Pubblicazione: (2025)
PINA: Prompt Injection Attack against Navigation Agents
di: Liu, Jiani, et al.
Pubblicazione: (2026)
di: Liu, Jiani, et al.
Pubblicazione: (2026)
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?
di: Pedro, Rodrigo, et al.
Pubblicazione: (2023)
di: Pedro, Rodrigo, et al.
Pubblicazione: (2023)
Bypassing Prompt Injection Detectors through Evasive Injections
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026)
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
AgentWatcher: A Rule-based Prompt Injection Monitor
di: Wang, Yanting, et al.
Pubblicazione: (2026)
di: Wang, Yanting, et al.
Pubblicazione: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
di: Chen, Yulin, et al.
Pubblicazione: (2024)
di: Chen, Yulin, et al.
Pubblicazione: (2024)
Cybersecurity AI: Hacking the AI Hackers via Prompt Injection
di: Mayoral-Vilches, Víctor, et al.
Pubblicazione: (2025)
di: Mayoral-Vilches, Víctor, et al.
Pubblicazione: (2025)
StruQ: Defending Against Prompt Injection with Structured Queries
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Strengthening Polymorphic Prompt Assembling: Dynamic Separator Generation Against Emerging Prompt Injection Attacks
di: Dorzhiev, Nima, et al.
Pubblicazione: (2026)
di: Dorzhiev, Nima, et al.
Pubblicazione: (2026)
Defeating Prompt Injections by Design
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2025)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2025)
Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives
di: Khodayari, Soheil, et al.
Pubblicazione: (2026)
di: Khodayari, Soheil, et al.
Pubblicazione: (2026)
Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2025)
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2025)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
di: An, Bang, et al.
Pubblicazione: (2024)
di: An, Bang, et al.
Pubblicazione: (2024)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems
di: Alam, Md Takrim Ul, et al.
Pubblicazione: (2026)
di: Alam, Md Takrim Ul, et al.
Pubblicazione: (2026)
Fingerprinting LLMs via Prompt Injection
di: Hu, Yuepeng, et al.
Pubblicazione: (2025)
di: Hu, Yuepeng, et al.
Pubblicazione: (2025)
PromptArmor: Simple yet Effective Prompt Injection Defenses
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
di: Wang, Jackson
Pubblicazione: (2026)
di: Wang, Jackson
Pubblicazione: (2026)
Documenti analoghi
-
Breaking the Prompt Wall (I): A Real-World Case Study of Attacking ChatGPT via Lightweight Prompt Injection
di: Chang, Xiangyu, et al.
Pubblicazione: (2025) -
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024) -
Privacy Artifact ConnecTor (PACT): Embedding Enterprise Artifacts for Compliance AI Agents
di: Fang, Chenhao, et al.
Pubblicazione: (2025) -
Evaluation of Prompt Injection Defenses in Large Language Models
di: Deep, Priyal, et al.
Pubblicazione: (2026) -
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals
di: Akinrele, Akindoyin, et al.
Pubblicazione: (2026)