Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiawen, Gupta, Pritha, Habernal, Ivan, Hüllermeier, Eyke
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910957533396992
author Wang, Jiawen
Gupta, Pritha
Habernal, Ivan
Hüllermeier, Eyke
author_facet Wang, Jiawen
Gupta, Pritha
Habernal, Ivan
Hüllermeier, Eyke
contents Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to different prompt-based attacks, generating harmful content or sensitive information. Both closed-source and open-source LLMs are underinvestigated for these attacks. This paper studies effective prompt injection attacks against the $\mathbf{14}$ most popular open-source LLMs on five attack benchmarks. Current metrics only consider successful attacks, whereas our proposed Attack Success Probability (ASP) also captures uncertainty in the model's response, reflecting ambiguity in attack feasibility. By comprehensively analyzing the effectiveness of prompt injection attacks, we propose a simple and effective hypnotism attack; results show that this attack causes aligned language models, including Stablelm2, Mistral, Openchat, and Vicuna, to generate objectionable behaviors, achieving around $90$% ASP. They also indicate that our ignore prefix attacks can break all $\mathbf{14}$ open-source LLMs, achieving over $60$% ASP on a multi-categorical dataset. We find that moderately well-known LLMs exhibit higher vulnerability to prompt injection attacks, highlighting the need to raise public awareness and prioritize efficient mitigation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14368
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
Wang, Jiawen
Gupta, Pritha
Habernal, Ivan
Hüllermeier, Eyke
Cryptography and Security
Computation and Language
Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to different prompt-based attacks, generating harmful content or sensitive information. Both closed-source and open-source LLMs are underinvestigated for these attacks. This paper studies effective prompt injection attacks against the $\mathbf{14}$ most popular open-source LLMs on five attack benchmarks. Current metrics only consider successful attacks, whereas our proposed Attack Success Probability (ASP) also captures uncertainty in the model's response, reflecting ambiguity in attack feasibility. By comprehensively analyzing the effectiveness of prompt injection attacks, we propose a simple and effective hypnotism attack; results show that this attack causes aligned language models, including Stablelm2, Mistral, Openchat, and Vicuna, to generate objectionable behaviors, achieving around $90$% ASP. They also indicate that our ignore prefix attacks can break all $\mathbf{14}$ open-source LLMs, achieving over $60$% ASP on a multi-categorical dataset. We find that moderately well-known LLMs exhibit higher vulnerability to prompt injection attacks, highlighting the need to raise public awareness and prioritize efficient mitigation strategies.
title Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2505.14368