Goal-guided Generative Prompt Injection Attack on Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Chong, Jin, Mingyu, Yu, Qinkai, Liu, Chengzhi, Xue, Haochen, Jin, Xiaobo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
di: Yu, Ye, et al.
Pubblicazione: (2026)
di: Yu, Ye, et al.
Pubblicazione: (2026)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
Assessing Prompt Injection Risks in 200+ Custom GPTs
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
di: Xiong, Junjie, et al.
Pubblicazione: (2025)
di: Xiong, Junjie, et al.
Pubblicazione: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
di: Zhang, Chong, et al.
Pubblicazione: (2025)
di: Zhang, Chong, et al.
Pubblicazione: (2025)
Prompt Injection Attacks on Large Language Models in Oncology
di: Clusmann, Jan, et al.
Pubblicazione: (2024)
di: Clusmann, Jan, et al.
Pubblicazione: (2024)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
di: Li, Zongze, et al.
Pubblicazione: (2025)
di: Li, Zongze, et al.
Pubblicazione: (2025)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
di: Guo, Ruoqi, et al.
Pubblicazione: (2026)
di: Guo, Ruoqi, et al.
Pubblicazione: (2026)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
DMFI: A Dual-Modality Log Analysis Framework for Insider Threat Detection with LoRA-Tuned Language Models
di: Kong, Kaichuan, et al.
Pubblicazione: (2025)
di: Kong, Kaichuan, et al.
Pubblicazione: (2025)
Prompt Injection as Role Confusion
di: Ye, Charles, et al.
Pubblicazione: (2026)
di: Ye, Charles, et al.
Pubblicazione: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
di: Deep, Priyal, et al.
Pubblicazione: (2026)
di: Deep, Priyal, et al.
Pubblicazione: (2026)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
di: Wang, Libo
Pubblicazione: (2024)
di: Wang, Libo
Pubblicazione: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
di: Sun, Xiongtao, et al.
Pubblicazione: (2024)
di: Sun, Xiongtao, et al.
Pubblicazione: (2024)
Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2024)
di: Shu, Dong, et al.
Pubblicazione: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
di: Yu, Miao, et al.
Pubblicazione: (2024)
di: Yu, Miao, et al.
Pubblicazione: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
di: Cao, Yuanpu, et al.
Pubblicazione: (2023)
di: Cao, Yuanpu, et al.
Pubblicazione: (2023)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
di: Wang, Haozhen, et al.
Pubblicazione: (2026)
di: Wang, Haozhen, et al.
Pubblicazione: (2026)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
di: Tu, Shangqing, et al.
Pubblicazione: (2024)
di: Tu, Shangqing, et al.
Pubblicazione: (2024)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
di: Park, Junyoung, et al.
Pubblicazione: (2026)
di: Park, Junyoung, et al.
Pubblicazione: (2026)
May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
di: Pandya, Nishit V., et al.
Pubblicazione: (2025)
di: Pandya, Nishit V., et al.
Pubblicazione: (2025)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
di: Zhou, Guanghao, et al.
Pubblicazione: (2025)
di: Zhou, Guanghao, et al.
Pubblicazione: (2025)
Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection
di: Wang, Zhilong, et al.
Pubblicazione: (2024)
di: Wang, Zhilong, et al.
Pubblicazione: (2024)
Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models
di: Ganiuly, Daniyal, et al.
Pubblicazione: (2025)
di: Ganiuly, Daniyal, et al.
Pubblicazione: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
Prompt Injection attack against LLM-integrated Applications
di: Liu, Yi, et al.
Pubblicazione: (2023)
di: Liu, Yi, et al.
Pubblicazione: (2023)
An Investigation on Group Query Hallucination Attacks
di: Miao, Kehao, et al.
Pubblicazione: (2025)
di: Miao, Kehao, et al.
Pubblicazione: (2025)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
di: Xu, Zihao, et al.
Pubblicazione: (2024)
di: Xu, Zihao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
di: Yu, Ye, et al.
Pubblicazione: (2026) -
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026) -
Assessing Prompt Injection Risks in 200+ Custom GPTs
di: Yu, Jiahao, et al.
Pubblicazione: (2023) -
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
di: Chae, Kyubyung, et al.
Pubblicazione: (2025) -
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)