JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Weidi, Ma, Siyuan, Liu, Xiaogeng, Guo, Xiaoyu, Xiao, Chaowei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
ShallowJail: Steering Jailbreaks against Large Language Models
von: Liu, Shang, et al.
Veröffentlicht: (2026)
von: Liu, Shang, et al.
Veröffentlicht: (2026)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
von: Nian, Yi, et al.
Veröffentlicht: (2025)
von: Nian, Yi, et al.
Veröffentlicht: (2025)
Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
OET: Optimization-based prompt injection Evaluation Toolkit
von: Pan, Jinsheng, et al.
Veröffentlicht: (2025)
von: Pan, Jinsheng, et al.
Veröffentlicht: (2025)
MMA-Diffusion: MultiModal Attack on Diffusion Models
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models
von: Luo, Weidi, et al.
Veröffentlicht: (2026)
von: Luo, Weidi, et al.
Veröffentlicht: (2026)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
von: Bi, Ting, et al.
Veröffentlicht: (2025)
von: Bi, Ting, et al.
Veröffentlicht: (2025)
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
SoK: Robustness in Large Language Models against Jailbreak Attacks
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
von: Yeke, Doguhuan, et al.
Veröffentlicht: (2026)
von: Yeke, Doguhuan, et al.
Veröffentlicht: (2026)
LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states
von: Ibanez-Lissen, Luis, et al.
Veröffentlicht: (2024)
von: Ibanez-Lissen, Luis, et al.
Veröffentlicht: (2024)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
von: Li, Hongyi, et al.
Veröffentlicht: (2024)
von: Li, Hongyi, et al.
Veröffentlicht: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
Prototype-Guided Robust Learning against Backdoor Attacks
von: Guo, Wei, et al.
Veröffentlicht: (2025)
von: Guo, Wei, et al.
Veröffentlicht: (2025)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
von: Piet, Julien, et al.
Veröffentlicht: (2025)
von: Piet, Julien, et al.
Veröffentlicht: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2024)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
von: Ma, Siyuan, et al.
Veröffentlicht: (2024) -
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025) -
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
von: Wang, Peiran, et al.
Veröffentlicht: (2024) -
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026) -
ShallowJail: Steering Jailbreaks against Large Language Models
von: Liu, Shang, et al.
Veröffentlicht: (2026)