InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhan, Qiusi, Liang, Zhixiang, Ying, Zifan, Kang, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
por: An, Hengyu, et al.
Publicado: (2025)
por: An, Hengyu, et al.
Publicado: (2025)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026)
por: Wang, Peiran, et al.
Publicado: (2026)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
por: Liu, Yinuo, et al.
Publicado: (2025)
por: Liu, Yinuo, et al.
Publicado: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
por: Hines, Keegan, et al.
Publicado: (2024)
por: Hines, Keegan, et al.
Publicado: (2024)
AI Agents May Always Fall for Prompt Injections
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
por: He, Yu, et al.
Publicado: (2026)
por: He, Yu, et al.
Publicado: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
por: Jia, Feiran, et al.
Publicado: (2024)
por: Jia, Feiran, et al.
Publicado: (2024)
QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
por: Xie, Yuchong, et al.
Publicado: (2025)
por: Xie, Yuchong, et al.
Publicado: (2025)
An Early Categorization of Prompt Injection Attacks on Large Language Models
por: Rossi, Sippo, et al.
Publicado: (2024)
por: Rossi, Sippo, et al.
Publicado: (2024)
LLM Agents can Autonomously Hack Websites
por: Fang, Richard, et al.
Publicado: (2024)
por: Fang, Richard, et al.
Publicado: (2024)
Goal-guided Generative Prompt Injection Attack on Large Language Models
por: Zhang, Chong, et al.
Publicado: (2024)
por: Zhang, Chong, et al.
Publicado: (2024)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
por: Yan, Jun, et al.
Publicado: (2023)
por: Yan, Jun, et al.
Publicado: (2023)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
por: Wang, Xilong, et al.
Publicado: (2026)
por: Wang, Xilong, et al.
Publicado: (2026)
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
por: Liang, Zi, et al.
Publicado: (2024)
por: Liang, Zi, et al.
Publicado: (2024)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
por: Guo, Ruoqi, et al.
Publicado: (2026)
por: Guo, Ruoqi, et al.
Publicado: (2026)
Multi-Agent Collaboration in Incident Response with Large Language Models
por: Liu, Zefang
Publicado: (2024)
por: Liu, Zefang
Publicado: (2024)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
por: Zhao, Lei, et al.
Publicado: (2026)
por: Zhao, Lei, et al.
Publicado: (2026)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
por: Sahoo, Devanshu, et al.
Publicado: (2025)
por: Sahoo, Devanshu, et al.
Publicado: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
por: Rassul, Yassin H., et al.
Publicado: (2026)
por: Rassul, Yassin H., et al.
Publicado: (2026)
Fingerprinting LLMs via Prompt Injection
por: Hu, Yuepeng, et al.
Publicado: (2025)
por: Hu, Yuepeng, et al.
Publicado: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
por: Yu, Zhiyuan, et al.
Publicado: (2024)
por: Yu, Zhiyuan, et al.
Publicado: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
por: Qiao, Yuxuan, et al.
Publicado: (2025)
por: Qiao, Yuxuan, et al.
Publicado: (2025)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
por: Yan, Lecheng, et al.
Publicado: (2026)
por: Yan, Lecheng, et al.
Publicado: (2026)
Defense Against Indirect Prompt Injection via Tool Result Parsing
por: Yu, Qiang, et al.
Publicado: (2026)
por: Yu, Qiang, et al.
Publicado: (2026)
Prompt Stealing Attacks Against Large Language Models
por: Sha, Zeyang, et al.
Publicado: (2024)
por: Sha, Zeyang, et al.
Publicado: (2024)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
por: Kimura, Subaru, et al.
Publicado: (2024)
por: Kimura, Subaru, et al.
Publicado: (2024)
Prompt Injection Attack to Tool Selection in LLM Agents
por: Shi, Jiawen, et al.
Publicado: (2025)
por: Shi, Jiawen, et al.
Publicado: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
por: Xu, Huiyu, et al.
Publicado: (2024)
por: Xu, Huiyu, et al.
Publicado: (2024)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
por: Wang, Jiawen, et al.
Publicado: (2025)
por: Wang, Jiawen, et al.
Publicado: (2025)
Separator Injection Attack: Uncovering Dialogue Biases in Large Language Models Caused by Role Separators
por: Li, Xitao, et al.
Publicado: (2025)
por: Li, Xitao, et al.
Publicado: (2025)
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
por: Alizadeh, Meysam, et al.
Publicado: (2025)
por: Alizadeh, Meysam, et al.
Publicado: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
por: Xu, Chejian, et al.
Publicado: (2024)
por: Xu, Chejian, et al.
Publicado: (2024)
Overthinking Loops in Agents: A Structural Risk via MCP Tools
por: Lee, Yohan, et al.
Publicado: (2026)
por: Lee, Yohan, et al.
Publicado: (2026)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
por: Gomaa, Amr, et al.
Publicado: (2025)
por: Gomaa, Amr, et al.
Publicado: (2025)
Ejemplares similares
-
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025) -
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024) -
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
por: An, Hengyu, et al.
Publicado: (2025) -
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026) -
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
por: Liu, Yinuo, et al.
Publicado: (2025)