How Not to Detect Prompt Injections with an LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Choudhary, Sarthak, Anshumaan, Divyam, Palumbo, Nils, Jha, Somesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dependency-Aware Privacy for Multi-turn Agents
by: Anshumaan, Divyam, et al.
Published: (2026)
by: Anshumaan, Divyam, et al.
Published: (2026)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
by: Wang, Zi, et al.
Published: (2024)
by: Wang, Zi, et al.
Published: (2024)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026)
by: Choudhary, Sarthak, et al.
Published: (2026)
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
Pr$εε$mpt: Sanitizing Sensitive Prompts for LLMs
by: Chowdhury, Amrita Roy, et al.
Published: (2025)
by: Chowdhury, Amrita Roy, et al.
Published: (2025)
Formal Policy Enforcement for Real-World Agentic Systems
by: Palumbo, Nils, et al.
Published: (2026)
by: Palumbo, Nils, et al.
Published: (2026)
Attacking Byzantine Robust Aggregation in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2023)
by: Choudhary, Sarthak, et al.
Published: (2023)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
by: Maiorano, Alexandre Cristovão
Published: (2026)
by: Maiorano, Alexandre Cristovão
Published: (2026)
Prompt Injection Attacks on Large Language Models in Oncology
by: Clusmann, Jan, et al.
Published: (2024)
by: Clusmann, Jan, et al.
Published: (2024)
Trust No AI: Prompt Injection Along The CIA Security Triad
by: Rehberger, Johann
Published: (2024)
by: Rehberger, Johann
Published: (2024)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
by: Benjamin, Victoria, et al.
Published: (2024)
by: Benjamin, Victoria, et al.
Published: (2024)
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
by: Liu, Ting-Chun, et al.
Published: (2025)
by: Liu, Ting-Chun, et al.
Published: (2025)
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Agent Security is a Systems Problem
by: Christodorescu, Mihai, et al.
Published: (2026)
by: Christodorescu, Mihai, et al.
Published: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
by: Zhang, Mohan, et al.
Published: (2026)
by: Zhang, Mohan, et al.
Published: (2026)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
by: Hackett, William, et al.
Published: (2025)
by: Hackett, William, et al.
Published: (2025)
Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems
by: Atta, Hammad, et al.
Published: (2025)
by: Atta, Hammad, et al.
Published: (2025)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
by: Liu, Mingrui, et al.
Published: (2026)
by: Liu, Mingrui, et al.
Published: (2026)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
by: Jia, Feiran, et al.
Published: (2024)
by: Jia, Feiran, et al.
Published: (2024)
PIArena: A Platform for Prompt Injection Evaluation
by: Geng, Runpeng, et al.
Published: (2026)
by: Geng, Runpeng, et al.
Published: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
by: Liu, Yupei, et al.
Published: (2023)
by: Liu, Yupei, et al.
Published: (2023)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
by: Correia, Pedro H. Barcha, et al.
Published: (2026)
by: Correia, Pedro H. Barcha, et al.
Published: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
by: Geng, Runpeng, et al.
Published: (2025)
by: Geng, Runpeng, et al.
Published: (2025)
Efficient Malware Detection with Optimized Learning on High-Dimensional Features
by: Choudhary, Aditya, et al.
Published: (2025)
by: Choudhary, Aditya, et al.
Published: (2025)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
by: Aghaei, Ehsan, et al.
Published: (2025)
by: Aghaei, Ehsan, et al.
Published: (2025)
False Data Injection Attack Detection in Edge-based Smart Metering Networks with Federated Learning
by: Uddin, Md Raihan, et al.
Published: (2024)
by: Uddin, Md Raihan, et al.
Published: (2024)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
by: Kharma, Mohammed, et al.
Published: (2026)
by: Kharma, Mohammed, et al.
Published: (2026)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Generating Synthetic Health Sensor Data for Privacy-Preserving Wearable Stress Detection
by: Lange, Lucas, et al.
Published: (2024)
by: Lange, Lucas, et al.
Published: (2024)
Detection of False Data Injection Attacks (FDIA) on Power Dynamical Systems With a State Prediction Method
by: Sahu, Abhijeet, et al.
Published: (2024)
by: Sahu, Abhijeet, et al.
Published: (2024)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
How Catastrophic is Your LLM? Certifying Risk in Conversation
by: Wang, Chengxiao, et al.
Published: (2025)
by: Wang, Chengxiao, et al.
Published: (2025)
SoK: Watermarking for AI-Generated Content
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
by: Yang, Xiaoxue, et al.
Published: (2025)
by: Yang, Xiaoxue, et al.
Published: (2025)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
by: Zeng, Qirun, et al.
Published: (2025)
by: Zeng, Qirun, et al.
Published: (2025)
Similar Items
-
Dependency-Aware Privacy for Multi-turn Agents
by: Anshumaan, Divyam, et al.
Published: (2026) -
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
by: Wang, Zi, et al.
Published: (2024) -
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026) -
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
by: Choudhary, Sarthak, et al.
Published: (2025) -
Pr$εε$mpt: Sanitizing Sensitive Prompts for LLMs
by: Chowdhury, Amrita Roy, et al.
Published: (2025)