A Critical Evaluation of Defenses against Prompt Injection Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Jia, Yuqi, Shao, Zedian, Liu, Yupei, Jia, Jinyuan, Song, Dawn, Gong, Neil Zhenqiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
por: Liu, Yupei, et al.
Publicado: (2023)
por: Liu, Yupei, et al.
Publicado: (2023)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
por: Shao, Zedian, et al.
Publicado: (2024)
por: Shao, Zedian, et al.
Publicado: (2024)
Evaluating LLM-based Personal Information Extraction and Countermeasures
por: Liu, Yupei, et al.
Publicado: (2024)
por: Liu, Yupei, et al.
Publicado: (2024)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
por: Zhang, Mohan, et al.
Publicado: (2026)
por: Zhang, Mohan, et al.
Publicado: (2026)
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
por: Shao, Zedian, et al.
Publicado: (2026)
por: Shao, Zedian, et al.
Publicado: (2026)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
por: Liu, Yinuo, et al.
Publicado: (2025)
por: Liu, Yinuo, et al.
Publicado: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
por: Wang, Reachal, et al.
Publicado: (2025)
por: Wang, Reachal, et al.
Publicado: (2025)
PromptArmor: Simple yet Effective Prompt Injection Defenses
por: Shi, Tianneng, et al.
Publicado: (2025)
por: Shi, Tianneng, et al.
Publicado: (2025)
Refusing Safe Prompts for Multi-modal Large Language Models
por: Shao, Zedian, et al.
Publicado: (2024)
por: Shao, Zedian, et al.
Publicado: (2024)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
por: Shi, Jiawen, et al.
Publicado: (2024)
por: Shi, Jiawen, et al.
Publicado: (2024)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
por: Wang, Xilong, et al.
Publicado: (2026)
por: Wang, Xilong, et al.
Publicado: (2026)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
por: Zou, Wei, et al.
Publicado: (2025)
por: Zou, Wei, et al.
Publicado: (2025)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
por: Yeo, Andrew, et al.
Publicado: (2025)
por: Yeo, Andrew, et al.
Publicado: (2025)
PIArena: A Platform for Prompt Injection Evaluation
por: Geng, Runpeng, et al.
Publicado: (2026)
por: Geng, Runpeng, et al.
Publicado: (2026)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
por: Zhao, Kaixiang, et al.
Publicado: (2025)
por: Zhao, Kaixiang, et al.
Publicado: (2025)
A Survey of Model Extraction Attacks and Defenses in Distributed Computing Environments
por: Zhao, Kaixiang, et al.
Publicado: (2025)
por: Zhao, Kaixiang, et al.
Publicado: (2025)
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
por: Zhao, Kaixiang, et al.
Publicado: (2025)
por: Zhao, Kaixiang, et al.
Publicado: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
por: Deep, Priyal, et al.
Publicado: (2026)
por: Deep, Priyal, et al.
Publicado: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2026)
por: Jia, Yuqi, et al.
Publicado: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
por: Zhu, Kaijie, et al.
Publicado: (2025)
por: Zhu, Kaijie, et al.
Publicado: (2025)
TrojFM: Resource-efficient Backdoor Attacks against Very Large Foundation Models
por: Nie, Yuzhou., et al.
Publicado: (2024)
por: Nie, Yuzhou., et al.
Publicado: (2024)
Competitive Advantage Attacks to Decentralized Federated Learning
por: Jia, Yuqi, et al.
Publicado: (2023)
por: Jia, Yuqi, et al.
Publicado: (2023)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
por: Bhatt, Manish, et al.
Publicado: (2026)
por: Bhatt, Manish, et al.
Publicado: (2026)
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
por: Kim, Juhee, et al.
Publicado: (2026)
por: Kim, Juhee, et al.
Publicado: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
por: Wang, Zhun, et al.
Publicado: (2025)
por: Wang, Zhun, et al.
Publicado: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
por: Xiang, Chong, et al.
Publicado: (2026)
por: Xiang, Chong, et al.
Publicado: (2026)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
por: Yang, Yuchen, et al.
Publicado: (2024)
por: Yang, Yuchen, et al.
Publicado: (2024)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
por: Hui, Bo, et al.
Publicado: (2024)
por: Hui, Bo, et al.
Publicado: (2024)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
por: Li, Yanjie, et al.
Publicado: (2025)
por: Li, Yanjie, et al.
Publicado: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
por: Cao, Tri, et al.
Publicado: (2026)
por: Cao, Tri, et al.
Publicado: (2026)
MalTool: Malicious Tool Attacks on LLM Agents
por: Hu, Yuepeng, et al.
Publicado: (2026)
por: Hu, Yuepeng, et al.
Publicado: (2026)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
por: Panebianco, Francesco, et al.
Publicado: (2025)
por: Panebianco, Francesco, et al.
Publicado: (2025)
Prompt Injection Attack to Tool Selection in LLM Agents
por: Shi, Jiawen, et al.
Publicado: (2025)
por: Shi, Jiawen, et al.
Publicado: (2025)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
por: Truong, Vu Tuan, et al.
Publicado: (2026)
por: Truong, Vu Tuan, et al.
Publicado: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
por: Geng, Runpeng, et al.
Publicado: (2025)
por: Geng, Runpeng, et al.
Publicado: (2025)
Ejemplares similares
-
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025) -
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
por: Liu, Yupei, et al.
Publicado: (2025) -
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
por: Liu, Yupei, et al.
Publicado: (2023) -
SecInfer: Preventing Prompt Injection via Inference-time Scaling
por: Liu, Yupei, et al.
Publicado: (2025) -
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
por: Shao, Zedian, et al.
Publicado: (2024)