Attention Tracker: Detecting Prompt Injection Attacks in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Hung, Kuo-Han, Ko, Ching-Yun, Rawat, Ambrish, Chung, I-Hsin, Hsu, Winston H., Chen, Pin-Yu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025)
por: Zizzo, Giulio, et al.
Publicado: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
por: Zhong, Yinan, et al.
Publicado: (2025)
por: Zhong, Yinan, et al.
Publicado: (2025)
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
por: Yin, Yu, et al.
Publicado: (2026)
por: Yin, Yu, et al.
Publicado: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2026)
por: Jia, Yuqi, et al.
Publicado: (2026)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
por: Cui, Yu, et al.
Publicado: (2025)
por: Cui, Yu, et al.
Publicado: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
por: Jaiswal, Piyush, et al.
Publicado: (2026)
por: Jaiswal, Piyush, et al.
Publicado: (2026)
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
por: Liu, Ting-Chun, et al.
Publicado: (2025)
por: Liu, Ting-Chun, et al.
Publicado: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
por: Yeo, Andrew, et al.
Publicado: (2025)
por: Yeo, Andrew, et al.
Publicado: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
por: Wang, Jiawen, et al.
Publicado: (2025)
por: Wang, Jiawen, et al.
Publicado: (2025)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection
por: Koide, Takashi, et al.
Publicado: (2026)
por: Koide, Takashi, et al.
Publicado: (2026)
Prompt Injection Attack to Tool Selection in LLM Agents
por: Shi, Jiawen, et al.
Publicado: (2025)
por: Shi, Jiawen, et al.
Publicado: (2025)
PINA: Prompt Injection Attack against Navigation Agents
por: Liu, Jiani, et al.
Publicado: (2026)
por: Liu, Jiani, et al.
Publicado: (2026)
PromptLocate: Localizing Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
por: Hu, Xiaomeng, et al.
Publicado: (2025)
por: Hu, Xiaomeng, et al.
Publicado: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
por: Wang, Reachal, et al.
Publicado: (2025)
por: Wang, Reachal, et al.
Publicado: (2025)
MAD-Spear: A Conformity-Driven Prompt Injection Attack on Multi-Agent Debate Systems
por: Cui, Yu, et al.
Publicado: (2025)
por: Cui, Yu, et al.
Publicado: (2025)
Duumviri: Detecting Trackers and Mixed Trackers with a Breakage Detector
por: Shuang, He, et al.
Publicado: (2024)
por: Shuang, He, et al.
Publicado: (2024)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
por: Zou, Wei, et al.
Publicado: (2025)
por: Zou, Wei, et al.
Publicado: (2025)
From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?
por: Pedro, Rodrigo, et al.
Publicado: (2023)
por: Pedro, Rodrigo, et al.
Publicado: (2023)
Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization
por: Verma, Ishaan, et al.
Publicado: (2025)
por: Verma, Ishaan, et al.
Publicado: (2025)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
por: Yang, Yuchen, et al.
Publicado: (2024)
por: Yang, Yuchen, et al.
Publicado: (2024)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
por: Wang, Che, et al.
Publicado: (2026)
por: Wang, Che, et al.
Publicado: (2026)
Strengthening Polymorphic Prompt Assembling: Dynamic Separator Generation Against Emerging Prompt Injection Attacks
por: Dorzhiev, Nima, et al.
Publicado: (2026)
por: Dorzhiev, Nima, et al.
Publicado: (2026)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
por: Chen, Yulin, et al.
Publicado: (2026)
por: Chen, Yulin, et al.
Publicado: (2026)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
por: Wang, Mengxiao, et al.
Publicado: (2025)
por: Wang, Mengxiao, et al.
Publicado: (2025)
Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification
por: Jayathilaka, Hasini
Publicado: (2025)
por: Jayathilaka, Hasini
Publicado: (2025)
Fingerprinting LLMs via Prompt Injection
por: Hu, Yuepeng, et al.
Publicado: (2025)
por: Hu, Yuepeng, et al.
Publicado: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
por: Xiong, Chen, et al.
Publicado: (2026)
por: Xiong, Chen, et al.
Publicado: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
por: Wang, Jackson
Publicado: (2026)
por: Wang, Jackson
Publicado: (2026)
Securing AI Agents Against Prompt Injection Attacks
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
Ejemplares similares
-
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025) -
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
por: Zhong, Yinan, et al.
Publicado: (2025) -
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025) -
Can Indirect Prompt Injection Attacks Be Detected and Removed?
por: Chen, Yulin, et al.
Publicado: (2025) -
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)