The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detection
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Corll, J Alex |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection
von: Corll, J Alex
Veröffentlicht: (2026)
von: Corll, J Alex
Veröffentlicht: (2026)
Defeating Prompt Injections by Design
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2025)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
Defending against Indirect Prompt Injection by Instruction Detection
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers
von: KumarRavindran, Santhosh
Veröffentlicht: (2025)
von: KumarRavindran, Santhosh
Veröffentlicht: (2025)
Bypassing Prompt Injection Detectors through Evasive Injections
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
von: Du, Mengyao, et al.
Veröffentlicht: (2026)
von: Du, Mengyao, et al.
Veröffentlicht: (2026)
How Not to Detect Prompt Injections with an LLM
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
PromptArmor: Simple yet Effective Prompt Injection Defenses
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models
von: Ganiuly, Daniyal, et al.
Veröffentlicht: (2025)
von: Ganiuly, Daniyal, et al.
Veröffentlicht: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
von: Chia-Pei, et al.
Veröffentlicht: (2026)
von: Chia-Pei, et al.
Veröffentlicht: (2026)
CASCADE: A Cascaded Hybrid Defense Architecture for Prompt Injection Detection in MCP-Based Systems
von: Turgut, İpek Abasıkeleş, et al.
Veröffentlicht: (2026)
von: Turgut, İpek Abasıkeleş, et al.
Veröffentlicht: (2026)
F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents
von: Ren, Yupeng
Veröffentlicht: (2024)
von: Ren, Yupeng
Veröffentlicht: (2024)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
von: Dziemian, Mateusz, et al.
Veröffentlicht: (2026)
von: Dziemian, Mateusz, et al.
Veröffentlicht: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Prompt Injection 2.0: Hybrid AI Threats
von: McHugh, Jeremy, et al.
Veröffentlicht: (2025)
von: McHugh, Jeremy, et al.
Veröffentlicht: (2025)
Assessing Prompt Injection Risks in 200+ Custom GPTs
von: Yu, Jiahao, et al.
Veröffentlicht: (2023)
von: Yu, Jiahao, et al.
Veröffentlicht: (2023)
Securing AI Agents Against Prompt Injection Attacks
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
von: Li, Zongze, et al.
Veröffentlicht: (2025)
von: Li, Zongze, et al.
Veröffentlicht: (2025)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
von: Bhatt, Manish, et al.
Veröffentlicht: (2026)
von: Bhatt, Manish, et al.
Veröffentlicht: (2026)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
von: Yeo, Andrew, et al.
Veröffentlicht: (2025)
von: Yeo, Andrew, et al.
Veröffentlicht: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
CourtGuard: A Local, Multiagent Prompt Injection Classifier
von: Wu, Isaac, et al.
Veröffentlicht: (2025)
von: Wu, Isaac, et al.
Veröffentlicht: (2025)
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
von: Xiong, Junjie, et al.
Veröffentlicht: (2025)
von: Xiong, Junjie, et al.
Veröffentlicht: (2025)
Prompt Injection Attacks on Large Language Models in Oncology
von: Clusmann, Jan, et al.
Veröffentlicht: (2024)
von: Clusmann, Jan, et al.
Veröffentlicht: (2024)
Adversarial Reinforcement Learning for Detecting False Data Injection Attacks in Vehicular Routing
von: Eghtesad, Taha, et al.
Veröffentlicht: (2026)
von: Eghtesad, Taha, et al.
Veröffentlicht: (2026)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2025)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2025)
Enhancing SQL Injection Detection and Prevention Using Generative Models
von: Dasari, Naga Sai, et al.
Veröffentlicht: (2025)
von: Dasari, Naga Sai, et al.
Veröffentlicht: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
von: Suo, Xuchen
Veröffentlicht: (2024)
von: Suo, Xuchen
Veröffentlicht: (2024)
Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
von: Chang, Hongyan, et al.
Veröffentlicht: (2026)
von: Chang, Hongyan, et al.
Veröffentlicht: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
von: Cao, Tri, et al.
Veröffentlicht: (2025)
von: Cao, Tri, et al.
Veröffentlicht: (2025)
PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection
von: Corll, J Alex
Veröffentlicht: (2026) -
Defeating Prompt Injections by Design
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2025) -
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025) -
Defending against Indirect Prompt Injection by Instruction Detection
von: Wen, Tongyu, et al.
Veröffentlicht: (2025) -
SecInfer: Preventing Prompt Injection via Inference-time Scaling
von: Liu, Yupei, et al.
Veröffentlicht: (2025)