Detecting Prompt Injection Attacks Against Application Using Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Shaheer, Safwan, Islam, G. M. Refatul, Hamid, Mohammad Rafid, Khan, Md. Abrar Faiaz, Faruk, Md. Omar, Nur, Yaseen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
Subjective Question Generation and Answer Evaluation using NLP
by: Islam, G. M. Refatul, et al.
Published: (2025)
by: Islam, G. M. Refatul, et al.
Published: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Attacking interpretable NLP systems
by: Abdukhamidov, Eldor, et al.
Published: (2025)
by: Abdukhamidov, Eldor, et al.
Published: (2025)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications
by: Motlagh, Farzad Nourmohammadzadeh, et al.
Published: (2026)
by: Motlagh, Farzad Nourmohammadzadeh, et al.
Published: (2026)
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
by: Hackett, William, et al.
Published: (2025)
by: Hackett, William, et al.
Published: (2025)
Guardians of the Web: The Evolution and Future of Website Information Security
by: Islam, Md Saiful, et al.
Published: (2025)
by: Islam, Md Saiful, et al.
Published: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
by: Adebimpe, Adetayo, et al.
Published: (2025)
by: Adebimpe, Adetayo, et al.
Published: (2025)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
Quantifying Return on Security Controls in LLM Systems
by: Moulton, Richard Helder, et al.
Published: (2025)
by: Moulton, Richard Helder, et al.
Published: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
by: Gowda, Ishrith
Published: (2026)
by: Gowda, Ishrith
Published: (2026)
Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
by: Kurian, Ashley, et al.
Published: (2025)
by: Kurian, Ashley, et al.
Published: (2025)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)
by: Ge, Yuxu
Published: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025)
by: Grofsky, Matthew
Published: (2025)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
by: Palaniappan, Nikitha M., et al.
Published: (2026)
by: Palaniappan, Nikitha M., et al.
Published: (2026)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
by: Bommarito II, Michael J.
Published: (2025)
by: Bommarito II, Michael J.
Published: (2025)
Towards Agentic Investigation of Security Alerts
by: Eilertsen, Even, et al.
Published: (2026)
by: Eilertsen, Even, et al.
Published: (2026)
Influence of Autoencoder Latent Space on Classifying IoT CoAP Attacks
by: García-Ordás, María Teresa, et al.
Published: (2026)
by: García-Ordás, María Teresa, et al.
Published: (2026)
From Attack Descriptions to Vulnerabilities: A Sentence Transformer-Based Approach
by: Othman, Refat, et al.
Published: (2025)
by: Othman, Refat, et al.
Published: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
by: Wei, Xingyuan, et al.
Published: (2024)
by: Wei, Xingyuan, et al.
Published: (2024)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
by: Pai, Aaditya
Published: (2026)
by: Pai, Aaditya
Published: (2026)
DP-2Stage: Adapting Language Models as Differentially Private Tabular Data Generators
by: Afonja, Tejumade, et al.
Published: (2024)
by: Afonja, Tejumade, et al.
Published: (2024)
Enhancing Wide-Angle Image Using Narrow-Angle View of the Same Scene
by: Safwan, Hussain Md., et al.
Published: (2025)
by: Safwan, Hussain Md., et al.
Published: (2025)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
by: Wang, Haochuan Kevin, et al.
Published: (2026)
by: Wang, Haochuan Kevin, et al.
Published: (2026)
The Quantum State Continuity Problem and Temporal Enforcement Against Fork Attacks
by: Ünsal, Samet
Published: (2025)
by: Ünsal, Samet
Published: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
On Explaining with Attention Matrices
by: Naim, Omar, et al.
Published: (2024)
by: Naim, Omar, et al.
Published: (2024)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
Defending against Backdoor Attacks via Module Switching
by: Li, Weijun, et al.
Published: (2025)
by: Li, Weijun, et al.
Published: (2025)
Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs
by: Bakker, Isabelle, et al.
Published: (2025)
by: Bakker, Isabelle, et al.
Published: (2025)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
ConfusionPrompt: Practical Private Inference for Online Large Language Models
by: Mai, Peihua, et al.
Published: (2023)
by: Mai, Peihua, et al.
Published: (2023)
Similar Items
-
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
by: Shaheer, Safwan, et al.
Published: (2025) -
Subjective Question Generation and Answer Evaluation using NLP
by: Islam, G. M. Refatul, et al.
Published: (2025) -
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026) -
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026) -
Attacking interpretable NLP systems
by: Abdukhamidov, Eldor, et al.
Published: (2025)