Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Xingwei, Fok, Kar Wai, Thing, Vrizlynn L. L. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Network Attack Traffic Detection With Hybrid Quantum-Enhanced Convolution Neural Network
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
by: Yang, Zhuochen, et al.
Published: (2025)
by: Yang, Zhuochen, et al.
Published: (2025)
Enhancing Network Intrusion Detection Performance using Generative Adversarial Networks
by: Zhao, Xinxing, et al.
Published: (2024)
by: Zhao, Xinxing, et al.
Published: (2024)
Enhanced Consistency Bi-directional GAN (CBiGAN) for Malware Anomaly Detection
by: Wijayasiri, Thesath, et al.
Published: (2025)
by: Wijayasiri, Thesath, et al.
Published: (2025)
ExpIDS: A Drift-adaptable Network Intrusion Detection System With Improved Explainability
by: Kumar, Ayush, et al.
Published: (2025)
by: Kumar, Ayush, et al.
Published: (2025)
Exploring Emerging Trends in 5G Malicious Traffic Analysis and Incremental Learning Intrusion Detection Strategies
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Privacy preserving layer partitioning for Deep Neural Network models
by: Rajasekar, Kishore, et al.
Published: (2024)
by: Rajasekar, Kishore, et al.
Published: (2024)
A Survey of Transaction Tracing Techniques for Blockchain Systems
by: Kumar, Ayush, et al.
Published: (2025)
by: Kumar, Ayush, et al.
Published: (2025)
Evaluating The Explainability of State-of-the-Art Deep Learning-based Network Intrusion Detection Systems
by: Kumar, Ayush, et al.
Published: (2024)
by: Kumar, Ayush, et al.
Published: (2024)
CPE-Identifier: Automated CPE identification and CVE summaries annotation with Deep Learning and NLP
by: Hu, Wanyu, et al.
Published: (2024)
by: Hu, Wanyu, et al.
Published: (2024)
Privacy-Preserving Intrusion Detection using Convolutional Neural Networks
by: Kodys, Martin, et al.
Published: (2024)
by: Kodys, Martin, et al.
Published: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
ICL-EVADER: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their Defenses
by: He, Ningyuan, et al.
Published: (2026)
by: He, Ningyuan, et al.
Published: (2026)
Can Drift-Adaptive Malware Detectors Be Made Robust? Attacks and Defenses Under White-Box and Black-Box Threats
by: Li, Adrian Shuai, et al.
Published: (2026)
by: Li, Adrian Shuai, et al.
Published: (2026)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
by: Cheng, Ruoxi, et al.
Published: (2024)
by: Cheng, Ruoxi, et al.
Published: (2024)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
by: Yang, Yiqi, et al.
Published: (2024)
by: Yang, Yiqi, et al.
Published: (2024)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
by: Hong, Hanbin, et al.
Published: (2023)
by: Hong, Hanbin, et al.
Published: (2023)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
BDFirewall: Towards Effective and Expeditiously Black-Box Backdoor Defense in MLaaS
by: Li, Ye, et al.
Published: (2025)
by: Li, Ye, et al.
Published: (2025)
Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
by: Ennaji, Sabrine, et al.
Published: (2025)
by: Ennaji, Sabrine, et al.
Published: (2025)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
by: Lin, Xingwei, et al.
Published: (2026)
by: Lin, Xingwei, et al.
Published: (2026)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
$\textit{MMJ-Bench}$: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
by: Weng, Fenghua, et al.
Published: (2024)
by: Weng, Fenghua, et al.
Published: (2024)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
by: Zhou, Yuqi, et al.
Published: (2024)
by: Zhou, Yuqi, et al.
Published: (2024)
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
by: Zhang, Junke, et al.
Published: (2026)
by: Zhang, Junke, et al.
Published: (2026)
Similar Items
-
DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
by: Wang, Zihao, et al.
Published: (2025) -
Network Attack Traffic Detection With Hybrid Quantum-Enhanced Convolution Neural Network
by: Wang, Zihao, et al.
Published: (2025) -
CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
by: Yang, Zhuochen, et al.
Published: (2025) -
Enhancing Network Intrusion Detection Performance using Generative Adversarial Networks
by: Zhao, Xinxing, et al.
Published: (2024) -
Enhanced Consistency Bi-directional GAN (CBiGAN) for Malware Anomaly Detection
by: Wijayasiri, Thesath, et al.
Published: (2025)