DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zihao, Fok, Kar Wai, Thing, Vrizlynn L. L. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025)
by: Zhong, Xingwei, et al.
Published: (2025)
CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
by: Yang, Zhuochen, et al.
Published: (2025)
by: Yang, Zhuochen, et al.
Published: (2025)
Network Attack Traffic Detection With Hybrid Quantum-Enhanced Convolution Neural Network
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Exploring Emerging Trends in 5G Malicious Traffic Analysis and Incremental Learning Intrusion Detection Strategies
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
ExpIDS: A Drift-adaptable Network Intrusion Detection System With Improved Explainability
by: Kumar, Ayush, et al.
Published: (2025)
by: Kumar, Ayush, et al.
Published: (2025)
Enhancing Network Intrusion Detection Performance using Generative Adversarial Networks
by: Zhao, Xinxing, et al.
Published: (2024)
by: Zhao, Xinxing, et al.
Published: (2024)
Enhanced Consistency Bi-directional GAN (CBiGAN) for Malware Anomaly Detection
by: Wijayasiri, Thesath, et al.
Published: (2025)
by: Wijayasiri, Thesath, et al.
Published: (2025)
Privacy preserving layer partitioning for Deep Neural Network models
by: Rajasekar, Kishore, et al.
Published: (2024)
by: Rajasekar, Kishore, et al.
Published: (2024)
A Survey of Transaction Tracing Techniques for Blockchain Systems
by: Kumar, Ayush, et al.
Published: (2025)
by: Kumar, Ayush, et al.
Published: (2025)
Evaluating The Explainability of State-of-the-Art Deep Learning-based Network Intrusion Detection Systems
by: Kumar, Ayush, et al.
Published: (2024)
by: Kumar, Ayush, et al.
Published: (2024)
CPE-Identifier: Automated CPE identification and CVE summaries annotation with Deep Learning and NLP
by: Hu, Wanyu, et al.
Published: (2024)
by: Hu, Wanyu, et al.
Published: (2024)
Privacy-Preserving Intrusion Detection using Convolutional Neural Networks
by: Kodys, Martin, et al.
Published: (2024)
by: Kodys, Martin, et al.
Published: (2024)
Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
by: Mishra, Rina, et al.
Published: (2025)
by: Mishra, Rina, et al.
Published: (2025)
Magnitude-based Neuron Pruning for Backdoor Defens
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
To See or Not to See: A Privacy Threat Model for Digital Forensics in Crime Investigation
by: Raciti, Mario, et al.
Published: (2025)
by: Raciti, Mario, et al.
Published: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
by: Liang, Siyuan, et al.
Published: (2025)
by: Liang, Siyuan, et al.
Published: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
by: Hossain, S M Asif, et al.
Published: (2025)
by: Hossain, S M Asif, et al.
Published: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025)
by: Geng, Jianing, et al.
Published: (2025)
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
by: Derya, Kemal, et al.
Published: (2026)
by: Derya, Kemal, et al.
Published: (2026)
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
by: Chen, Luoyu, et al.
Published: (2026)
by: Chen, Luoyu, et al.
Published: (2026)
Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models
by: Wang, Yanting, et al.
Published: (2024)
by: Wang, Yanting, et al.
Published: (2024)
$\textit{MMJ-Bench}$: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
by: Weng, Fenghua, et al.
Published: (2024)
by: Weng, Fenghua, et al.
Published: (2024)
A Federated Learning Approach for Multi-stage Threat Analysis in Advanced Persistent Threat Campaigns
by: Nelles, Florian, et al.
Published: (2024)
by: Nelles, Florian, et al.
Published: (2024)
Towards an AI-Enhanced Cyber Threat Intelligence Processing Pipeline
by: Alevizos, Lampis, et al.
Published: (2024)
by: Alevizos, Lampis, et al.
Published: (2024)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026)
by: Zhu, Zhenhao, et al.
Published: (2026)
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era?
by: Bhandarkar, Avanti, et al.
Published: (2024)
by: Bhandarkar, Avanti, et al.
Published: (2024)
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
by: Zeng, Xiyu, et al.
Published: (2025)
by: Zeng, Xiyu, et al.
Published: (2025)
ORCA - An Automated Threat Analysis Pipeline for O-RAN Continuous Development
by: Klement, Felix, et al.
Published: (2026)
by: Klement, Felix, et al.
Published: (2026)
Ensemble Defense System: A Hybrid IDS Approach for Effective Cyber Threat Detection
by: Alharbi, Sarah, et al.
Published: (2024)
by: Alharbi, Sarah, et al.
Published: (2024)
Similar Items
-
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025) -
CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
by: Yang, Zhuochen, et al.
Published: (2025) -
Network Attack Traffic Detection With Hybrid Quantum-Enhanced Convolution Neural Network
by: Wang, Zihao, et al.
Published: (2025) -
Exploring Emerging Trends in 5G Malicious Traffic Analysis and Incremental Learning Intrusion Detection Strategies
by: Wang, Zihao, et al.
Published: (2024) -
ExpIDS: A Drift-adaptable Network Intrusion Detection System With Improved Explainability
by: Kumar, Ayush, et al.
Published: (2025)