PromptArmor: Simple yet Effective Prompt Injection Defenses
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Tianneng, Zhu, Kaijie, Wang, Zhun, Jia, Yuqi, Cai, Will, Liang, Weida, Wang, Haonan, Alzahrani, Hend, Lu, Joshua, Kawaguchi, Kenji, Alomair, Basel, Zhao, Xuandong, Wang, William Yang, Gong, Neil, Guo, Wenbo, Song, Dawn |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
por: Wang, Zhun, et al.
Publicado: (2025)
por: Wang, Zhun, et al.
Publicado: (2025)
Can LLMs Ask Good Questions?
por: Zhang, Yueheng, et al.
Publicado: (2025)
por: Zhang, Yueheng, et al.
Publicado: (2025)
Defending Against Prompt Injection with DataFilter
por: Wang, Yizhu, et al.
Publicado: (2025)
por: Wang, Yizhu, et al.
Publicado: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
por: Piet, Julien, et al.
Publicado: (2023)
por: Piet, Julien, et al.
Publicado: (2023)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
por: Zhu, Kaijie, et al.
Publicado: (2025)
por: Zhu, Kaijie, et al.
Publicado: (2025)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
por: Wang, Xilong, et al.
Publicado: (2026)
por: Wang, Xilong, et al.
Publicado: (2026)
Preventing Prompt Injection with Type-Directed Privilege Separation
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
por: Liu, Yupei, et al.
Publicado: (2023)
por: Liu, Yupei, et al.
Publicado: (2023)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
por: Cai, Will, et al.
Publicado: (2025)
por: Cai, Will, et al.
Publicado: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
por: Wang, Reachal, et al.
Publicado: (2025)
por: Wang, Reachal, et al.
Publicado: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2026)
por: Jia, Yuqi, et al.
Publicado: (2026)
Progent: Securing AI Agents with Privilege Control
por: Shi, Tianneng, et al.
Publicado: (2025)
por: Shi, Tianneng, et al.
Publicado: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
por: Liu, Yinuo, et al.
Publicado: (2025)
por: Liu, Yinuo, et al.
Publicado: (2025)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
por: Wang, Peiran, et al.
Publicado: (2025)
por: Wang, Peiran, et al.
Publicado: (2025)
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
por: Kim, Minbeom, et al.
Publicado: (2026)
por: Kim, Minbeom, et al.
Publicado: (2026)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
por: Zhang, Mohan, et al.
Publicado: (2026)
por: Zhang, Mohan, et al.
Publicado: (2026)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
por: Wang, Zhun, et al.
Publicado: (2025)
por: Wang, Zhun, et al.
Publicado: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
por: Nie, Yuzhou, et al.
Publicado: (2024)
por: Nie, Yuzhou, et al.
Publicado: (2024)
Defending Against Prompt Injection With a Few DefensiveTokens
por: Chen, Sizhe, et al.
Publicado: (2025)
por: Chen, Sizhe, et al.
Publicado: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
por: Yin, Chenlong, et al.
Publicado: (2026)
por: Yin, Chenlong, et al.
Publicado: (2026)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
por: Bhatt, Manish, et al.
Publicado: (2026)
por: Bhatt, Manish, et al.
Publicado: (2026)
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
por: Kim, Juhee, et al.
Publicado: (2026)
por: Kim, Juhee, et al.
Publicado: (2026)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
por: Wang, Che, et al.
Publicado: (2026)
por: Wang, Che, et al.
Publicado: (2026)
FilterPrompt: A Simple yet Efficient Approach to Guide Image Appearance Transfer in Diffusion Models
por: Wang, Xi, et al.
Publicado: (2024)
por: Wang, Xi, et al.
Publicado: (2024)
Evaluation of Prompt Injection Defenses in Large Language Models
por: Deep, Priyal, et al.
Publicado: (2026)
por: Deep, Priyal, et al.
Publicado: (2026)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
por: Wang, Yihan, et al.
Publicado: (2025)
por: Wang, Yihan, et al.
Publicado: (2025)
Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model
por: Yin, Junhui, et al.
Publicado: (2026)
por: Yin, Junhui, et al.
Publicado: (2026)
Improving LLM Safety Alignment with Dual-Objective Optimization
por: Zhao, Xuandong, et al.
Publicado: (2025)
por: Zhao, Xuandong, et al.
Publicado: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
por: Potter, Yujin, et al.
Publicado: (2025)
por: Potter, Yujin, et al.
Publicado: (2025)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
por: Zhu, Kaijie, et al.
Publicado: (2023)
por: Zhu, Kaijie, et al.
Publicado: (2023)
Can LLMs Follow Simple Rules?
por: Mu, Norman, et al.
Publicado: (2023)
por: Mu, Norman, et al.
Publicado: (2023)
WebInject: Prompt Injection Attack to Web Agents
por: Wang, Xilong, et al.
Publicado: (2025)
por: Wang, Xilong, et al.
Publicado: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
por: Yeo, Andrew, et al.
Publicado: (2025)
por: Yeo, Andrew, et al.
Publicado: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
Ejemplares similares
-
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025) -
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
por: Wang, Zhun, et al.
Publicado: (2025) -
Can LLMs Ask Good Questions?
por: Zhang, Yueheng, et al.
Publicado: (2025) -
Defending Against Prompt Injection with DataFilter
por: Wang, Yizhu, et al.
Publicado: (2025) -
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
por: Piet, Julien, et al.
Publicado: (2023)