Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
Fuente:
arXiv
Saved in:
| Main Authors: | Pasquini, Dario, Kornaropoulos, Evgenios M., Ateniese, Giuseppe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMmap: Fingerprinting For Large Language Models
by: Pasquini, Dario, et al.
Published: (2024)
by: Pasquini, Dario, et al.
Published: (2024)
When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
by: Pasquini, Dario, et al.
Published: (2025)
by: Pasquini, Dario, et al.
Published: (2025)
TENNOR: Trustworthy Execution for Neural Networks through Obliviousness and Retrievals
by: Qu, Zifan, et al.
Published: (2026)
by: Qu, Zifan, et al.
Published: (2026)
Cybersecurity AI: Hacking the AI Hackers via Prompt Injection
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
by: Xiang, Chong, et al.
Published: (2026)
by: Xiang, Chong, et al.
Published: (2026)
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
by: Wang, Zhilong, et al.
Published: (2025)
by: Wang, Zhilong, et al.
Published: (2025)
Agent Privilege Separation in OpenClaw: A Structural Defense Against Prompt Injection
by: Cheng, Darren, et al.
Published: (2026)
by: Cheng, Darren, et al.
Published: (2026)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
Universal Neural-Cracking-Machines: Self-Configurable Password Models from Auxiliary Data
by: Pasquini, Dario, et al.
Published: (2023)
by: Pasquini, Dario, et al.
Published: (2023)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
by: Bhatt, Manish, et al.
Published: (2026)
by: Bhatt, Manish, et al.
Published: (2026)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
by: Lan, Qianlong, et al.
Published: (2026)
by: Lan, Qianlong, et al.
Published: (2026)
PromptArmor: Simple yet Effective Prompt Injection Defenses
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
by: Zhong, Peter Yong, et al.
Published: (2025)
by: Zhong, Peter Yong, et al.
Published: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
by: Deep, Priyal, et al.
Published: (2026)
by: Deep, Priyal, et al.
Published: (2026)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
by: An, Hengyu, et al.
Published: (2025)
by: An, Hengyu, et al.
Published: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
by: Suo, Xuchen
Published: (2024)
by: Suo, Xuchen
Published: (2024)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
by: Hasan, Md. Mehedi, et al.
Published: (2025)
by: Hasan, Md. Mehedi, et al.
Published: (2025)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
by: Yeo, Andrew, et al.
Published: (2025)
by: Yeo, Andrew, et al.
Published: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
Defenses Against Prompt Attacks Learn Surface Heuristics
by: Li, Shawn, et al.
Published: (2026)
by: Li, Shawn, et al.
Published: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
ESLD (External Surrogate Latent Defense): A Latent-Space Architecture for Faster, Stronger Prompt-Injection Defense
by: Narendra, Yash
Published: (2026)
by: Narendra, Yash
Published: (2026)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
by: Wang, Che, et al.
Published: (2026)
by: Wang, Che, et al.
Published: (2026)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
by: Correia, Pedro H. Barcha, et al.
Published: (2026)
by: Correia, Pedro H. Barcha, et al.
Published: (2026)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
by: Zhao, Wei, et al.
Published: (2026)
by: Zhao, Wei, et al.
Published: (2026)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026)
by: Chia-Pei, et al.
Published: (2026)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025)
by: Rodriguez, Mikel, et al.
Published: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
by: Shi, Jiawen, et al.
Published: (2024)
by: Shi, Jiawen, et al.
Published: (2024)
The Coding Limits of Robust Watermarking for Generative Models
by: Francati, Danilo, et al.
Published: (2025)
by: Francati, Danilo, et al.
Published: (2025)
CASCADE: A Cascaded Hybrid Defense Architecture for Prompt Injection Detection in MCP-Based Systems
by: Turgut, İpek Abasıkeleş, et al.
Published: (2026)
by: Turgut, İpek Abasıkeleş, et al.
Published: (2026)
Similar Items
-
LLMmap: Fingerprinting For Large Language Models
by: Pasquini, Dario, et al.
Published: (2024) -
When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
by: Pasquini, Dario, et al.
Published: (2025) -
TENNOR: Trustworthy Execution for Neural Networks through Obliviousness and Retrievals
by: Qu, Zifan, et al.
Published: (2026) -
Cybersecurity AI: Hacking the AI Hackers via Prompt Injection
by: Mayoral-Vilches, Víctor, et al.
Published: (2025) -
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)