Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking
Fuente:
arXiv
Salvato in:
| Autori principali: | Cohen, Stav, Bitton, Ron, Nassi, Ben |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
di: Cohen, Stav, et al.
Pubblicazione: (2024)
di: Cohen, Stav, et al.
Pubblicazione: (2024)
A Jailbroken GenAI Model Can Cause Substantial Harm: GenAI-powered Applications are Vulnerable to PromptWares
di: Cohen, Stav, et al.
Pubblicazione: (2024)
di: Cohen, Stav, et al.
Pubblicazione: (2024)
Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
di: Nassi, Ben, et al.
Pubblicazione: (2025)
di: Nassi, Ben, et al.
Pubblicazione: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
E-MIA: Exam-Style Black-Box Membership Inference Attacks against RAG Systems
di: Guan, Zelin, et al.
Pubblicazione: (2026)
di: Guan, Zelin, et al.
Pubblicazione: (2026)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
di: Park, Junyoung, et al.
Pubblicazione: (2026)
di: Park, Junyoung, et al.
Pubblicazione: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
di: Xu, Feiyue, et al.
Pubblicazione: (2026)
di: Xu, Feiyue, et al.
Pubblicazione: (2026)
LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
di: Happe, Andreas, et al.
Pubblicazione: (2023)
di: Happe, Andreas, et al.
Pubblicazione: (2023)
The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism
di: Brodt, Oleg, et al.
Pubblicazione: (2026)
di: Brodt, Oleg, et al.
Pubblicazione: (2026)
Untargeted Jailbreak Attack
di: Huang, Xinzhe, et al.
Pubblicazione: (2025)
di: Huang, Xinzhe, et al.
Pubblicazione: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
di: Wang, Youze, et al.
Pubblicazione: (2025)
di: Wang, Youze, et al.
Pubblicazione: (2025)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
di: Song, Weiming, et al.
Pubblicazione: (2026)
di: Song, Weiming, et al.
Pubblicazione: (2026)
Data-Free Model-Related Attacks: Unleashing the Potential of Generative AI
di: Ye, Dayong, et al.
Pubblicazione: (2025)
di: Ye, Dayong, et al.
Pubblicazione: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
di: Tong, Haibo, et al.
Pubblicazione: (2025)
di: Tong, Haibo, et al.
Pubblicazione: (2025)
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
di: Probst, Benjamin, et al.
Pubblicazione: (2026)
di: Probst, Benjamin, et al.
Pubblicazione: (2026)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)
di: Liu, Xiaoqun, et al.
Pubblicazione: (2024)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
di: Liu, Shuyuan, et al.
Pubblicazione: (2025)
di: Liu, Shuyuan, et al.
Pubblicazione: (2025)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
di: Guo, Yangyang, et al.
Pubblicazione: (2025)
di: Guo, Yangyang, et al.
Pubblicazione: (2025)
Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks
di: Al-Lawati, Ali, et al.
Pubblicazione: (2026)
di: Al-Lawati, Ali, et al.
Pubblicazione: (2026)
The Hidden Threat in Plain Text: Attacking RAG Data Loaders
di: Castagnaro, Alberto, et al.
Pubblicazione: (2025)
di: Castagnaro, Alberto, et al.
Pubblicazione: (2025)
EVA: Editing for Versatile Alignment against Jailbreaks
di: Wang, Yi, et al.
Pubblicazione: (2026)
di: Wang, Yi, et al.
Pubblicazione: (2026)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
di: Qi, Senmao, et al.
Pubblicazione: (2025)
di: Qi, Senmao, et al.
Pubblicazione: (2025)
FlipAttack: Jailbreak LLMs via Flipping
di: Liu, Yue, et al.
Pubblicazione: (2024)
di: Liu, Yue, et al.
Pubblicazione: (2024)
Synthetic Cancer -- Augmenting Worms with LLMs
di: Zimmerman, Benjamin, et al.
Pubblicazione: (2024)
di: Zimmerman, Benjamin, et al.
Pubblicazione: (2024)
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
di: Li, Yuying, et al.
Pubblicazione: (2024)
di: Li, Yuying, et al.
Pubblicazione: (2024)
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
di: Cunningham, Hoagy, et al.
Pubblicazione: (2026)
di: Cunningham, Hoagy, et al.
Pubblicazione: (2026)
Similarity-based Label Inference Attack against Training and Inference of Split Learning
di: Liu, Junlin, et al.
Pubblicazione: (2022)
di: Liu, Junlin, et al.
Pubblicazione: (2022)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
di: Fu, Haowei, et al.
Pubblicazione: (2025)
di: Fu, Haowei, et al.
Pubblicazione: (2025)
The System Prompt Is the Attack Surface: How LLM Agent Configuration Shapes Security and Creates Exploitable Vulnerabilities
di: Litvak, Ron
Pubblicazione: (2026)
di: Litvak, Ron
Pubblicazione: (2026)
ShallowJail: Steering Jailbreaks against Large Language Models
di: Liu, Shang, et al.
Pubblicazione: (2026)
di: Liu, Shang, et al.
Pubblicazione: (2026)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
di: Wang, Peiran, et al.
Pubblicazione: (2024)
di: Wang, Peiran, et al.
Pubblicazione: (2024)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
di: Shang, Zhengchun, et al.
Pubblicazione: (2025)
di: Shang, Zhengchun, et al.
Pubblicazione: (2025)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
di: Chen, Kejia, et al.
Pubblicazione: (2026)
di: Chen, Kejia, et al.
Pubblicazione: (2026)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
di: Zahran, Noureldin, et al.
Pubblicazione: (2025)
di: Zahran, Noureldin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
di: Cohen, Stav, et al.
Pubblicazione: (2024) -
A Jailbroken GenAI Model Can Cause Substantial Harm: GenAI-powered Applications are Vulnerable to PromptWares
di: Cohen, Stav, et al.
Pubblicazione: (2024) -
Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
di: Nassi, Ben, et al.
Pubblicazione: (2025) -
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024) -
E-MIA: Exam-Style Black-Box Membership Inference Attacks against RAG Systems
di: Guan, Zelin, et al.
Pubblicazione: (2026)