Securing AI Agents Against Prompt Injection Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramakrishnan, Badrinath, Balaji, Akshaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917093727797248
author Ramakrishnan, Badrinath
Balaji, Akshaya
author_facet Ramakrishnan, Badrinath
Balaji, Akshaya
contents Retrieval-augmented generation (RAG) systems have become widely used for enhancing large language model capabilities, but they introduce significant security vulnerabilities through prompt injection attacks. We present a comprehensive benchmark for evaluating prompt injection risks in RAG-enabled AI agents and propose a multi-layered defense framework. Our benchmark includes 847 adversarial test cases across five attack categories: direct injection, context manipulation, instruction override, data exfiltration, and cross-context contamination. We evaluate three defense mechanisms: content filtering with embedding-based anomaly detection, hierarchical system prompt guardrails, and multi-stage response verification, across seven state-of-the-art language models. Our combined framework reduces successful attack rates from 73.2% to 8.7% while maintaining 94.3% of baseline task performance. We release our benchmark dataset and defense implementation to support future research in AI agent security.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15759
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Securing AI Agents Against Prompt Injection Attacks
Ramakrishnan, Badrinath
Balaji, Akshaya
Cryptography and Security
Artificial Intelligence
Retrieval-augmented generation (RAG) systems have become widely used for enhancing large language model capabilities, but they introduce significant security vulnerabilities through prompt injection attacks. We present a comprehensive benchmark for evaluating prompt injection risks in RAG-enabled AI agents and propose a multi-layered defense framework. Our benchmark includes 847 adversarial test cases across five attack categories: direct injection, context manipulation, instruction override, data exfiltration, and cross-context contamination. We evaluate three defense mechanisms: content filtering with embedding-based anomaly detection, hierarchical system prompt guardrails, and multi-stage response verification, across seven state-of-the-art language models. Our combined framework reduces successful attack rates from 73.2% to 8.7% while maintaining 94.3% of baseline task performance. We release our benchmark dataset and defense implementation to support future research in AI agent security.
title Securing AI Agents Against Prompt Injection Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2511.15759