Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Yu, Sun, Sheng, Cheng, Shengjia, Liu, Teli, Li, Mingfeng, Liu, Min |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
por: Yan, Yu, et al.
Publicado: (2025)
por: Yan, Yu, et al.
Publicado: (2025)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
por: Wang, Xinkai, et al.
Publicado: (2025)
por: Wang, Xinkai, et al.
Publicado: (2025)
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
por: Mao, Xutao, et al.
Publicado: (2026)
por: Mao, Xutao, et al.
Publicado: (2026)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
por: Yin, Ziyi, et al.
Publicado: (2025)
por: Yin, Ziyi, et al.
Publicado: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
por: Li, Jie, et al.
Publicado: (2024)
por: Li, Jie, et al.
Publicado: (2024)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
por: You, Wenhao, et al.
Publicado: (2025)
por: You, Wenhao, et al.
Publicado: (2025)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
por: Dou, Yipu, et al.
Publicado: (2026)
por: Dou, Yipu, et al.
Publicado: (2026)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
por: Hao, Shuyang, et al.
Publicado: (2025)
por: Hao, Shuyang, et al.
Publicado: (2025)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
por: Wang, Youze, et al.
Publicado: (2025)
por: Wang, Youze, et al.
Publicado: (2025)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
por: Zou, Quanchen, et al.
Publicado: (2026)
por: Zou, Quanchen, et al.
Publicado: (2026)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
por: Liu, Mingrui, et al.
Publicado: (2025)
por: Liu, Mingrui, et al.
Publicado: (2025)
SoK: Robustness in Large Language Models against Jailbreak Attacks
por: Xu, Feiyue, et al.
Publicado: (2026)
por: Xu, Feiyue, et al.
Publicado: (2026)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
por: Teng, Ma, et al.
Publicado: (2024)
por: Teng, Ma, et al.
Publicado: (2024)
$\textit{MMJ-Bench}$: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
por: Weng, Fenghua, et al.
Publicado: (2024)
por: Weng, Fenghua, et al.
Publicado: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
por: Lu, Lin, et al.
Publicado: (2024)
por: Lu, Lin, et al.
Publicado: (2024)
DREAM: Dynamic Red-teaming across Environments for AI Models
por: Lu, Liming, et al.
Publicado: (2025)
por: Lu, Liming, et al.
Publicado: (2025)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
por: Yang, Yijun, et al.
Publicado: (2025)
por: Yang, Yijun, et al.
Publicado: (2025)
Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning
por: Wang, Zhaoqi, et al.
Publicado: (2026)
por: Wang, Zhaoqi, et al.
Publicado: (2026)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
por: Ma, Siyuan, et al.
Publicado: (2024)
por: Ma, Siyuan, et al.
Publicado: (2024)
Practical Reasoning Interruption Attacks on Reasoning Large Language Models
por: Cui, Yu, et al.
Publicado: (2025)
por: Cui, Yu, et al.
Publicado: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
por: Zhang, Chenyu, et al.
Publicado: (2025)
por: Zhang, Chenyu, et al.
Publicado: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
por: Tu, Shangqing, et al.
Publicado: (2024)
por: Tu, Shangqing, et al.
Publicado: (2024)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
por: Zeng, Xiyu, et al.
Publicado: (2025)
por: Zeng, Xiyu, et al.
Publicado: (2025)
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization
por: Liu, Aofan, et al.
Publicado: (2025)
por: Liu, Aofan, et al.
Publicado: (2025)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
por: Kumar, Divyanshu, et al.
Publicado: (2025)
por: Kumar, Divyanshu, et al.
Publicado: (2025)
FlipAttack: Jailbreak LLMs via Flipping
por: Liu, Yue, et al.
Publicado: (2024)
por: Liu, Yue, et al.
Publicado: (2024)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
por: Xu, Zihao, et al.
Publicado: (2024)
por: Xu, Zihao, et al.
Publicado: (2024)
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
por: Zhang, Deyue, et al.
Publicado: (2025)
por: Zhang, Deyue, et al.
Publicado: (2025)
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
por: Zhou, Yuqi, et al.
Publicado: (2024)
por: Zhou, Yuqi, et al.
Publicado: (2024)
Distract Large Language Models for Automatic Jailbreak Attack
por: Xiao, Zeguan, et al.
Publicado: (2024)
por: Xiao, Zeguan, et al.
Publicado: (2024)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
por: Yuan, Zenghui, et al.
Publicado: (2025)
por: Yuan, Zenghui, et al.
Publicado: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
por: Liu, Xiaogeng, et al.
Publicado: (2025)
por: Liu, Xiaogeng, et al.
Publicado: (2025)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
por: Xu, Zihao, et al.
Publicado: (2024)
por: Xu, Zihao, et al.
Publicado: (2024)
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
por: Huang, Xijie, et al.
Publicado: (2024)
por: Huang, Xijie, et al.
Publicado: (2024)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
por: Du, Yuhao, et al.
Publicado: (2024)
por: Du, Yuhao, et al.
Publicado: (2024)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
por: Piet, Julien, et al.
Publicado: (2025)
por: Piet, Julien, et al.
Publicado: (2025)
Ejemplares similares
-
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
por: Yan, Yu, et al.
Publicado: (2025) -
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
por: Wang, Xinkai, et al.
Publicado: (2025) -
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
por: Mao, Xutao, et al.
Publicado: (2026) -
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
por: Yin, Ziyi, et al.
Publicado: (2025) -
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
por: Li, Jie, et al.
Publicado: (2024)