Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Chong, Zagieboylo, Drew, Ghosh, Shaona, Kariyappa, Sanjay, Greshake, Kai, Xiao, Hanshen, Xiao, Chaowei, Suh, G. Edward
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915903372787712
author Xiang, Chong
Zagieboylo, Drew
Ghosh, Shaona
Kariyappa, Sanjay
Greshake, Kai
Xiao, Hanshen
Xiao, Chaowei
Suh, G. Edward
author_facet Xiang, Chong
Zagieboylo, Drew
Ghosh, Shaona
Kariyappa, Sanjay
Greshake, Kai
Xiao, Hanshen
Xiao, Chaowei
Suh, G. Edward
contents AI agents, predominantly powered by large language models (LLMs), are vulnerable to indirect prompt injection, in which malicious instructions embedded in untrusted data can trigger dangerous agent actions. This position paper discusses our vision for system-level defenses against indirect prompt injection attacks. We articulate three positions: (1) dynamic replanning and security policy updates are often necessary for dynamic tasks and realistic environments; (2) certain context-dependent security decisions would still require LLMs (or other learned models), but should only be made within system designs that strictly constrain what the model can observe and decide; (3) in inherently ambiguous cases, personalization and human interaction should be treated as core design considerations. In addition to our main positions, we discuss limitations of existing benchmarks that can create a false sense of utility and security. We also highlight the value of system-level defenses, which serve as the skeleton of agentic systems by structuring and controlling agent behaviors, integrating rule-based and model-based security checks, and enabling more targeted research on model robustness and human interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2603_30016
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
Xiang, Chong
Zagieboylo, Drew
Ghosh, Shaona
Kariyappa, Sanjay
Greshake, Kai
Xiao, Hanshen
Xiao, Chaowei
Suh, G. Edward
Cryptography and Security
Artificial Intelligence
AI agents, predominantly powered by large language models (LLMs), are vulnerable to indirect prompt injection, in which malicious instructions embedded in untrusted data can trigger dangerous agent actions. This position paper discusses our vision for system-level defenses against indirect prompt injection attacks. We articulate three positions: (1) dynamic replanning and security policy updates are often necessary for dynamic tasks and realistic environments; (2) certain context-dependent security decisions would still require LLMs (or other learned models), but should only be made within system designs that strictly constrain what the model can observe and decide; (3) in inherently ambiguous cases, personalization and human interaction should be treated as core design considerations. In addition to our main positions, we discuss limitations of existing benchmarks that can create a false sense of utility and security. We also highlight the value of system-level defenses, which serve as the skeleton of agentic systems by structuring and controlling agent behaviors, integrating rule-based and model-based security checks, and enabling more targeted research on model robustness and human interaction.
title Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2603.30016