Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Zimo, Wang, Xunguang, Li, Zongjie, Ma, Pingchuan, Gao, Yudong, Wu, Daoyuan, Yan, Xincheng, Tian, Tian, Wang, Shuai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909913425379328
author Ji, Zimo
Wang, Xunguang
Li, Zongjie
Ma, Pingchuan
Gao, Yudong
Wu, Daoyuan
Yan, Xincheng
Tian, Tian
Wang, Shuai
author_facet Ji, Zimo
Wang, Xunguang
Li, Zongjie
Ma, Pingchuan
Gao, Yudong
Wu, Daoyuan
Yan, Xincheng
Tian, Tian
Wang, Shuai
contents Large Language Model (LLM)-based agents with function-calling capabilities are increasingly deployed, but remain vulnerable to Indirect Prompt Injection (IPI) attacks that hijack their tool calls. In response, numerous IPI-centric defense frameworks have emerged. However, these defenses are fragmented, lacking a unified taxonomy and comprehensive evaluation. In this Systematization of Knowledge (SoK), we present the first comprehensive analysis of IPI-centric defense frameworks. We introduce a comprehensive taxonomy of these defenses, classifying them along five dimensions. We then thoroughly assess the security and usability of representative defense frameworks. Through analysis of defensive failures in the assessment, we identify six root causes of defense circumvention. Based on these findings, we design three novel adaptive attacks that significantly improve attack success rates targeting specific frameworks, demonstrating the severity of the flaws in these defenses. Our paper provides a foundation and critical insights for the future development of more secure and usable IPI-centric agent defense frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15203
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
Ji, Zimo
Wang, Xunguang
Li, Zongjie
Ma, Pingchuan
Gao, Yudong
Wu, Daoyuan
Yan, Xincheng
Tian, Tian
Wang, Shuai
Cryptography and Security
Artificial Intelligence
Large Language Model (LLM)-based agents with function-calling capabilities are increasingly deployed, but remain vulnerable to Indirect Prompt Injection (IPI) attacks that hijack their tool calls. In response, numerous IPI-centric defense frameworks have emerged. However, these defenses are fragmented, lacking a unified taxonomy and comprehensive evaluation. In this Systematization of Knowledge (SoK), we present the first comprehensive analysis of IPI-centric defense frameworks. We introduce a comprehensive taxonomy of these defenses, classifying them along five dimensions. We then thoroughly assess the security and usability of representative defense frameworks. Through analysis of defensive failures in the assessment, we identify six root causes of defense circumvention. Based on these findings, we design three novel adaptive attacks that significantly improve attack success rates targeting specific frameworks, demonstrating the severity of the flaws in these defenses. Our paper provides a foundation and critical insights for the future development of more secure and usable IPI-centric agent defense frameworks.
title Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2511.15203