Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shaheer, Safwan, Islam, G. M. Refatul, Hamid, Mohammad Rafid, Jilan, Tahsin Zaman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915684423827456
author Shaheer, Safwan
Islam, G. M. Refatul
Hamid, Mohammad Rafid
Jilan, Tahsin Zaman
author_facet Shaheer, Safwan
Islam, G. M. Refatul
Hamid, Mohammad Rafid
Jilan, Tahsin Zaman
contents In this fast-evolving area of LLMs, our paper discusses the significant security risk presented by prompt injection attacks. It focuses on small open-sourced models, specifically the LLaMA family of models. We introduce novel defense mechanisms capable of generating automatic defenses and systematically evaluate said generated defenses against a comprehensive set of benchmarked attacks. Thus, we empirically demonstrated the improvement proposed by our approach in mitigating goal-hijacking vulnerabilities in LLMs. Our work recognizes the increasing relevance of small open-sourced LLMs and their potential for broad deployments on edge devices, aligning with future trends in LLM applications. We contribute to the greater ecosystem of open-source LLMs and their security in the following: (1) assessing present prompt-based defenses against the latest attacks, (2) introducing a new framework using a seed defense (Chain Of Thoughts) to refine the defense prompts iteratively, and (3) showing significant improvements in detecting goal hijacking attacks. Out strategies significantly reduce the success rates of the attacks and false detection rates while at the same time effectively detecting goal-hijacking capabilities, paving the way for more secure and efficient deployments of small and open-source LLMs in resource-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
Shaheer, Safwan
Islam, G. M. Refatul
Hamid, Mohammad Rafid
Jilan, Tahsin Zaman
Cryptography and Security
Artificial Intelligence
D.4.6; I.2.7
In this fast-evolving area of LLMs, our paper discusses the significant security risk presented by prompt injection attacks. It focuses on small open-sourced models, specifically the LLaMA family of models. We introduce novel defense mechanisms capable of generating automatic defenses and systematically evaluate said generated defenses against a comprehensive set of benchmarked attacks. Thus, we empirically demonstrated the improvement proposed by our approach in mitigating goal-hijacking vulnerabilities in LLMs. Our work recognizes the increasing relevance of small open-sourced LLMs and their potential for broad deployments on edge devices, aligning with future trends in LLM applications. We contribute to the greater ecosystem of open-source LLMs and their security in the following: (1) assessing present prompt-based defenses against the latest attacks, (2) introducing a new framework using a seed defense (Chain Of Thoughts) to refine the defense prompts iteratively, and (3) showing significant improvements in detecting goal hijacking attacks. Out strategies significantly reduce the success rates of the attacks and false detection rates while at the same time effectively detecting goal-hijacking capabilities, paving the way for more secure and efficient deployments of small and open-source LLMs in resource-constrained environments.
title Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
topic Cryptography and Security
Artificial Intelligence
D.4.6; I.2.7
url https://arxiv.org/abs/2512.16307