From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Siyuan, Long, Zhuohan, Fan, Zhihao, Wei, Zhongyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Playing Language Game with LLMs Leads to Jailbreaking
von: Peng, Yu, et al.
Veröffentlicht: (2024)
von: Peng, Yu, et al.
Veröffentlicht: (2024)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
von: Wang, Yiqi, et al.
Veröffentlicht: (2024)
von: Wang, Yiqi, et al.
Veröffentlicht: (2024)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
Poisoned LangChain: Jailbreak LLMs by LangChain
von: Wang, Ziqiu, et al.
Veröffentlicht: (2024)
von: Wang, Ziqiu, et al.
Veröffentlicht: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
von: Shen, Guangyu, et al.
Veröffentlicht: (2024)
von: Shen, Guangyu, et al.
Veröffentlicht: (2024)
Defending LLMs against Jailbreaking Attacks via Backtranslation
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
von: Jiang, Shixin, et al.
Veröffentlicht: (2024)
von: Jiang, Shixin, et al.
Veröffentlicht: (2024)
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
Jailbreaking to Jailbreak
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
von: Noughabi, Havva Alizadeh, et al.
Veröffentlicht: (2025)
von: Noughabi, Havva Alizadeh, et al.
Veröffentlicht: (2025)
ALaRM: Align Language Models via Hierarchical Rewards Modeling
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
von: Weng, Zixuan, et al.
Veröffentlicht: (2025)
von: Weng, Zixuan, et al.
Veröffentlicht: (2025)
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
von: Zhang, Guibin, et al.
Veröffentlicht: (2025)
von: Zhang, Guibin, et al.
Veröffentlicht: (2025)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
von: Luo, Yifan, et al.
Veröffentlicht: (2024)
von: Luo, Yifan, et al.
Veröffentlicht: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
The Cost of Thinking: Increased Jailbreak Risk in Large Language Models
von: Yang, Fan
Veröffentlicht: (2025)
von: Yang, Fan
Veröffentlicht: (2025)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
von: Chen, Sirry, et al.
Veröffentlicht: (2026)
von: Chen, Sirry, et al.
Veröffentlicht: (2026)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
von: Long, Zhuohan, et al.
Veröffentlicht: (2026)
von: Long, Zhuohan, et al.
Veröffentlicht: (2026)
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
von: Choukrani, Omar, et al.
Veröffentlicht: (2025)
von: Choukrani, Omar, et al.
Veröffentlicht: (2025)
Affordance Benchmark for MLLMs
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Defending against Jailbreak through Early Exit Generation of Large Language Models
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024)
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
von: Kim, Mirae, et al.
Veröffentlicht: (2026)
von: Kim, Mirae, et al.
Veröffentlicht: (2026)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
Beyond LLMs: Advancing the Landscape of Complex Reasoning
von: Chu-Carroll, Jennifer, et al.
Veröffentlicht: (2024)
von: Chu-Carroll, Jennifer, et al.
Veröffentlicht: (2024)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models
von: Nguyen, Tri, et al.
Veröffentlicht: (2025)
von: Nguyen, Tri, et al.
Veröffentlicht: (2025)
Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025) -
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
von: Wang, Siyuan, et al.
Veröffentlicht: (2024) -
Playing Language Game with LLMs Leads to Jailbreaking
von: Peng, Yu, et al.
Veröffentlicht: (2024) -
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024) -
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
von: Wang, Yiqi, et al.
Veröffentlicht: (2024)