Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
Fuente:
arXiv
Salvato in:
| Autori principali: | Ahmed, Mohamed, Abdelmouty, Mohamed, Kim, Mingyu, Kandula, Gunvanth, Park, Alex, Davis, James C. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
di: Nihal, Ragib Amin, et al.
Pubblicazione: (2025)
di: Nihal, Ragib Amin, et al.
Pubblicazione: (2025)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
di: Li, Yucheng, et al.
Pubblicazione: (2025)
di: Li, Yucheng, et al.
Pubblicazione: (2025)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
di: Luo, Xuan, et al.
Pubblicazione: (2025)
di: Luo, Xuan, et al.
Pubblicazione: (2025)
BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
di: Nakka, Kalyan, et al.
Pubblicazione: (2025)
di: Nakka, Kalyan, et al.
Pubblicazione: (2025)
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
di: Yang, Guangyu, et al.
Pubblicazione: (2025)
di: Yang, Guangyu, et al.
Pubblicazione: (2025)
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
di: Tshimula, Jean Marie, et al.
Pubblicazione: (2024)
di: Tshimula, Jean Marie, et al.
Pubblicazione: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
Proactive defense against LLM Jailbreak
di: Zhao, Weiliang, et al.
Pubblicazione: (2025)
di: Zhao, Weiliang, et al.
Pubblicazione: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
di: Gu, Haoran, et al.
Pubblicazione: (2026)
di: Gu, Haoran, et al.
Pubblicazione: (2026)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
A Comprehensive Analysis of Routing Vulnerabilities and Defense Strategies in IoT Networks
di: Jae-Dong, Kim
Pubblicazione: (2024)
di: Jae-Dong, Kim
Pubblicazione: (2024)
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
di: Ren, Xiaoxue, et al.
Pubblicazione: (2025)
di: Ren, Xiaoxue, et al.
Pubblicazione: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
di: Pathade, Chetan
Pubblicazione: (2025)
di: Pathade, Chetan
Pubblicazione: (2025)
Local Frames: Exploiting Inherited Origins to Bypass Content Blockers
di: Ukani, Alisha, et al.
Pubblicazione: (2025)
di: Ukani, Alisha, et al.
Pubblicazione: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
di: Long, Zhuohang, et al.
Pubblicazione: (2025)
di: Long, Zhuohang, et al.
Pubblicazione: (2025)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
di: Afane, Mohamed, et al.
Pubblicazione: (2025)
di: Afane, Mohamed, et al.
Pubblicazione: (2025)
SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs
di: Shaaban, Mohamed, et al.
Pubblicazione: (2026)
di: Shaaban, Mohamed, et al.
Pubblicazione: (2026)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
di: Wu, Yu-Hang, et al.
Pubblicazione: (2025)
di: Wu, Yu-Hang, et al.
Pubblicazione: (2025)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
di: Gibbs, Tom, et al.
Pubblicazione: (2024)
di: Gibbs, Tom, et al.
Pubblicazione: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
di: Shen, Guobin, et al.
Pubblicazione: (2025)
di: Shen, Guobin, et al.
Pubblicazione: (2025)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
di: Li, Jindong, et al.
Pubblicazione: (2026)
di: Li, Jindong, et al.
Pubblicazione: (2026)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
di: Zhang, Chiyu, et al.
Pubblicazione: (2026)
di: Zhang, Chiyu, et al.
Pubblicazione: (2026)
Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities
di: Mouzouni, Charafeddine
Pubblicazione: (2026)
di: Mouzouni, Charafeddine
Pubblicazione: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
di: Huang, Ruixuan, et al.
Pubblicazione: (2025)
di: Huang, Ruixuan, et al.
Pubblicazione: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
di: Ouyang, Yang, et al.
Pubblicazione: (2025)
di: Ouyang, Yang, et al.
Pubblicazione: (2025)
LLM Jailbreak Detection for (Almost) Free!
di: Chen, Guorui, et al.
Pubblicazione: (2025)
di: Chen, Guorui, et al.
Pubblicazione: (2025)
Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
di: Polyakov, Alex, et al.
Pubblicazione: (2026)
di: Polyakov, Alex, et al.
Pubblicazione: (2026)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
di: Liu, Songyang, et al.
Pubblicazione: (2026)
di: Liu, Songyang, et al.
Pubblicazione: (2026)
Comprehensive Digital Forensics and Risk Mitigation Strategy for Modern Enterprises
di: Shaffi, Shamnad Mohamed
Pubblicazione: (2025)
di: Shaffi, Shamnad Mohamed
Pubblicazione: (2025)
Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Security
di: Kim, Hajun, et al.
Pubblicazione: (2025)
di: Kim, Hajun, et al.
Pubblicazione: (2025)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
di: Zhou, Yukai, et al.
Pubblicazione: (2025)
di: Zhou, Yukai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
di: Zeng, Yifan, et al.
Pubblicazione: (2024) -
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
di: Nihal, Ragib Amin, et al.
Pubblicazione: (2025) -
Testing the Limits of Jailbreaking Defenses with the Purple Problem
di: Kim, Taeyoun, et al.
Pubblicazione: (2024) -
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
di: Chen, Bocheng, et al.
Pubblicazione: (2024) -
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
di: Li, Yucheng, et al.
Pubblicazione: (2025)