WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xilong, Liu, Yinuo, Wang, Zhun, Song, Dawn, Gong, Neil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
WebInject: Prompt Injection Attack to Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2025)
di: Wang, Xilong, et al.
Pubblicazione: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
di: Shi, Jiawen, et al.
Pubblicazione: (2024)
di: Shi, Jiawen, et al.
Pubblicazione: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
di: Wang, Zhun, et al.
Pubblicazione: (2025)
di: Wang, Zhun, et al.
Pubblicazione: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
di: Zhang, Chong, et al.
Pubblicazione: (2024)
di: Zhang, Chong, et al.
Pubblicazione: (2024)
PromptArmor: Simple yet Effective Prompt Injection Defenses
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
di: Du, Mengyao, et al.
Pubblicazione: (2026)
di: Du, Mengyao, et al.
Pubblicazione: (2026)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
di: Syros, Georgios, et al.
Pubblicazione: (2026)
di: Syros, Georgios, et al.
Pubblicazione: (2026)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
di: Johnson, Sam, et al.
Pubblicazione: (2025)
di: Johnson, Sam, et al.
Pubblicazione: (2025)
A Framework for Formalizing LLM Agent Security
di: Siu, Vincent, et al.
Pubblicazione: (2026)
di: Siu, Vincent, et al.
Pubblicazione: (2026)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
di: Fujinuma, Yoshinari, et al.
Pubblicazione: (2026)
di: Fujinuma, Yoshinari, et al.
Pubblicazione: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
Prompt Injection as Role Confusion
di: Ye, Charles, et al.
Pubblicazione: (2026)
di: Ye, Charles, et al.
Pubblicazione: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
di: Wang, Xunguang, et al.
Pubblicazione: (2025)
di: Wang, Xunguang, et al.
Pubblicazione: (2025)
Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization
di: Verma, Ishaan, et al.
Pubblicazione: (2025)
di: Verma, Ishaan, et al.
Pubblicazione: (2025)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
di: Guo, Ruoqi, et al.
Pubblicazione: (2026)
di: Guo, Ruoqi, et al.
Pubblicazione: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
di: Gong, Bofan, et al.
Pubblicazione: (2025)
di: Gong, Bofan, et al.
Pubblicazione: (2025)
Prompt Injection attack against LLM-integrated Applications
di: Liu, Yi, et al.
Pubblicazione: (2023)
di: Liu, Yi, et al.
Pubblicazione: (2023)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
di: Nawal, Aditya, et al.
Pubblicazione: (2026)
di: Nawal, Aditya, et al.
Pubblicazione: (2026)
Large Language Model Sentinel: LLM Agent for Adversarial Purification
di: Lin, Guang, et al.
Pubblicazione: (2024)
di: Lin, Guang, et al.
Pubblicazione: (2024)
May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
di: Pandya, Nishit V., et al.
Pubblicazione: (2025)
di: Pandya, Nishit V., et al.
Pubblicazione: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
di: An, Hengyu, et al.
Pubblicazione: (2025)
di: An, Hengyu, et al.
Pubblicazione: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
Activation-Guided Local Editing for Jailbreaking Attacks
di: Wang, Jiecong, et al.
Pubblicazione: (2025)
di: Wang, Jiecong, et al.
Pubblicazione: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
di: Li, Xirui, et al.
Pubblicazione: (2024)
di: Li, Xirui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025) -
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
di: Liu, Yupei, et al.
Pubblicazione: (2025) -
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026) -
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2026) -
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)