Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
Fuente:
arXiv
Salvato in:
| Autori principali: | Jaiswal, Piyush, Pratap, Aaditya, Saraswati, Shreyansh, Kasyap, Harsh, Tripathy, Somanath |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fairness-Constrained Optimization Attack in Federated Learning
di: Kasyap, Harsh, et al.
Pubblicazione: (2025)
di: Kasyap, Harsh, et al.
Pubblicazione: (2025)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
di: Yeo, Andrew, et al.
Pubblicazione: (2025)
di: Yeo, Andrew, et al.
Pubblicazione: (2025)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
di: Guo, Yangyang, et al.
Pubblicazione: (2025)
di: Guo, Yangyang, et al.
Pubblicazione: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
di: Tong, Haibo, et al.
Pubblicazione: (2025)
di: Tong, Haibo, et al.
Pubblicazione: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
di: Suo, Xuchen
Pubblicazione: (2024)
di: Suo, Xuchen
Pubblicazione: (2024)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
di: Zahran, Noureldin, et al.
Pubblicazione: (2025)
di: Zahran, Noureldin, et al.
Pubblicazione: (2025)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
di: Park, Junyoung, et al.
Pubblicazione: (2026)
di: Park, Junyoung, et al.
Pubblicazione: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
di: Syros, Georgios, et al.
Pubblicazione: (2026)
di: Syros, Georgios, et al.
Pubblicazione: (2026)
FlipAttack: Jailbreak LLMs via Flipping
di: Liu, Yue, et al.
Pubblicazione: (2024)
di: Liu, Yue, et al.
Pubblicazione: (2024)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
di: Wang, Che, et al.
Pubblicazione: (2026)
di: Wang, Che, et al.
Pubblicazione: (2026)
Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization
di: Verma, Ishaan, et al.
Pubblicazione: (2025)
di: Verma, Ishaan, et al.
Pubblicazione: (2025)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
di: Hung, Kuo-Han, et al.
Pubblicazione: (2024)
di: Hung, Kuo-Han, et al.
Pubblicazione: (2024)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
di: Xiang, Chong, et al.
Pubblicazione: (2026)
di: Xiang, Chong, et al.
Pubblicazione: (2026)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
di: Tsmindashvili, Tatia, et al.
Pubblicazione: (2025)
di: Tsmindashvili, Tatia, et al.
Pubblicazione: (2025)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
di: Shang, Zhengchun, et al.
Pubblicazione: (2025)
di: Shang, Zhengchun, et al.
Pubblicazione: (2025)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
di: Cui, Tiehan, et al.
Pubblicazione: (2025)
di: Cui, Tiehan, et al.
Pubblicazione: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
di: Shi, Jiawen, et al.
Pubblicazione: (2024)
di: Shi, Jiawen, et al.
Pubblicazione: (2024)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
Untargeted Jailbreak Attack
di: Huang, Xinzhe, et al.
Pubblicazione: (2025)
di: Huang, Xinzhe, et al.
Pubblicazione: (2025)
Defenses Against Prompt Attacks Learn Surface Heuristics
di: Li, Shawn, et al.
Pubblicazione: (2026)
di: Li, Shawn, et al.
Pubblicazione: (2026)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
di: Cao, Tri, et al.
Pubblicazione: (2025)
di: Cao, Tri, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Fairness-Constrained Optimization Attack in Federated Learning
di: Kasyap, Harsh, et al.
Pubblicazione: (2025) -
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025) -
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
di: Zhang, Zheng, et al.
Pubblicazione: (2025) -
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025) -
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)