Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Zimo, Wang, Xunguang, Li, Zongjie, Ma, Pingchuan, Gao, Yudong, Wu, Daoyuan, Yan, Xincheng, Tian, Tian, Wang, Shuai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)
by: Huang, Ruixuan, et al.
Published: (2025)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026)
by: Wang, Xunguang, et al.
Published: (2026)
QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
by: Xie, Yuchong, et al.
Published: (2025)
by: Xie, Yuchong, et al.
Published: (2025)
Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
by: Li, Zongjie, et al.
Published: (2025)
by: Li, Zongjie, et al.
Published: (2025)
SEAL: Subspace-Anchored Watermarks for LLM Ownership
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
SkillReducer: Optimizing LLM Agent Skills for Token Efficiency
by: Gao, Yudong, et al.
Published: (2026)
by: Gao, Yudong, et al.
Published: (2026)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
Reframing LLM Agent Security as an Agent-Human Interaction Problem
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
KryptoPilot: An Open-World Knowledge-Augmented LLM Agent for Automated Cryptographic Exploitation
by: Liu, Xiaonan, et al.
Published: (2026)
by: Liu, Xiaonan, et al.
Published: (2026)
Disabling Self-Correction in Retrieval-Augmented Generation via Stealthy Retriever Poisoning
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
by: Wang, Yihan, et al.
Published: (2025)
by: Wang, Yihan, et al.
Published: (2025)
ZK-Value: A Practical Zero-Knowledge System for Verifiable Data Valuation
by: Wang, Zhaoyu, et al.
Published: (2026)
by: Wang, Zhaoyu, et al.
Published: (2026)
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
by: Wei, Qianshan, et al.
Published: (2025)
by: Wei, Qianshan, et al.
Published: (2025)
Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities
by: Mouzouni, Charafeddine
Published: (2026)
by: Mouzouni, Charafeddine
Published: (2026)
Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation
by: Andreucci, Biagio, et al.
Published: (2026)
by: Andreucci, Biagio, et al.
Published: (2026)
ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
by: Liu, Ruixuan, et al.
Published: (2024)
by: Liu, Ruixuan, et al.
Published: (2024)
Autonomous LLM Agent Worms: Cross-Platform Propagation, Automated Discovery and Temporal Re-Entry Defense
by: Zha, Mingming, et al.
Published: (2026)
by: Zha, Mingming, et al.
Published: (2026)
A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms
by: Acharya, Nirajan, et al.
Published: (2026)
by: Acharya, Nirajan, et al.
Published: (2026)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
by: Wu, Daoyuan, et al.
Published: (2024)
by: Wu, Daoyuan, et al.
Published: (2024)
Cyber-Physical Security Vulnerabilities Identification and Classification in Smart Manufacturing -- A Defense-in-Depth Driven Framework and Taxonomy
by: Rahman, Md Habibor, et al.
Published: (2024)
by: Rahman, Md Habibor, et al.
Published: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
by: Yang, Yixuan, et al.
Published: (2025)
by: Yang, Yixuan, et al.
Published: (2025)
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
by: Wang, Zeng, et al.
Published: (2026)
by: Wang, Zeng, et al.
Published: (2026)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
A Taxonomy of Attacks and Defenses in Split Learning
by: Shabbir, Aqsa, et al.
Published: (2025)
by: Shabbir, Aqsa, et al.
Published: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
SOK: A Taxonomy of Attack Vectors and Defense Strategies for Agentic Supply Chain Runtime
by: Jiang, Xiaochong, et al.
Published: (2026)
by: Jiang, Xiaochong, et al.
Published: (2026)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Differential Privacy for Network Assortativity
by: Ma, Fei, et al.
Published: (2025)
by: Ma, Fei, et al.
Published: (2025)
Diffusion-Guided Adversarial Perturbation Injection for Generalizable Defense Against Facial Manipulations
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models
by: Meng, Xiangtao, et al.
Published: (2026)
by: Meng, Xiangtao, et al.
Published: (2026)
Similar Items
-
Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework
by: Ji, Zimo, et al.
Published: (2026) -
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025) -
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025) -
SoK: Evaluating Jailbreak Guardrails for Large Language Models
by: Wang, Xunguang, et al.
Published: (2025) -
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)