Saved in:
Bibliographic Details
Main Authors: Zou, Wei, Liu, Yupei, Wang, Yanting, Chen, Ying, Gong, Neil, Jia, Jinyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.14005
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911395074801664
author Zou, Wei
Liu, Yupei
Wang, Yanting
Chen, Ying
Gong, Neil
Jia, Jinyuan
author_facet Zou, Wei
Liu, Yupei
Wang, Yanting
Chen, Ying
Gong, Neil
Jia, Jinyuan
contents LLM-integrated applications are vulnerable to prompt injection attacks, where an attacker contaminates the input to inject malicious instructions, causing the LLM to follow the attacker's intent instead of the original user's. Existing prompt injection detection methods often have sub-optimal performance and/or high computational overhead. In this work, we propose PIShield, an effective and efficient detection method based on the observation that instruction-tuned LLMs internally encode distinguishable signals for prompts containing injected instructions. PIShield leverages residual-stream representations and a simple linear classifier to detect prompt injection, without expensive model fine-tuning or response generation. We conduct extensive evaluations on a diverse set of short- and long-context benchmarks. The results show that PIShield consistently achieves low false positive and false negative rates, significantly outperforming existing baselines. These findings demonstrate that internal representations of instruction-tuned LLMs provide a powerful and practical foundation for prompt injection detection in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
Zou, Wei
Liu, Yupei
Wang, Yanting
Chen, Ying
Gong, Neil
Jia, Jinyuan
Cryptography and Security
Machine Learning
LLM-integrated applications are vulnerable to prompt injection attacks, where an attacker contaminates the input to inject malicious instructions, causing the LLM to follow the attacker's intent instead of the original user's. Existing prompt injection detection methods often have sub-optimal performance and/or high computational overhead. In this work, we propose PIShield, an effective and efficient detection method based on the observation that instruction-tuned LLMs internally encode distinguishable signals for prompts containing injected instructions. PIShield leverages residual-stream representations and a simple linear classifier to detect prompt injection, without expensive model fine-tuning or response generation. We conduct extensive evaluations on a diverse set of short- and long-context benchmarks. The results show that PIShield consistently achieves low false positive and false negative rates, significantly outperforming existing baselines. These findings demonstrate that internal representations of instruction-tuned LLMs provide a powerful and practical foundation for prompt injection detection in real-world applications.
title PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2510.14005