Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Wenzhuo, Wei, Zhipeng, Sun, Xiongtao, Ying, Zonghao, Zhang, Deyue, Yang, Dongdong, Zhang, Xiangzheng, Zou, Quanchen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
von: Chen, Moyang, et al.
Veröffentlicht: (2026)
von: Chen, Moyang, et al.
Veröffentlicht: (2026)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
von: Zou, Quanchen, et al.
Veröffentlicht: (2025)
von: Zou, Quanchen, et al.
Veröffentlicht: (2025)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
von: Zou, Quanchen, et al.
Veröffentlicht: (2026)
von: Zou, Quanchen, et al.
Veröffentlicht: (2026)
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
von: Zhang, Deyue, et al.
Veröffentlicht: (2025)
von: Zhang, Deyue, et al.
Veröffentlicht: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
von: Liu, Zhe, et al.
Veröffentlicht: (2026)
von: Liu, Zhe, et al.
Veröffentlicht: (2026)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Robust Privacy: Inference-Time Privacy through Certified Robustness
von: Jin, Jiankai, et al.
Veröffentlicht: (2026)
von: Jin, Jiankai, et al.
Veröffentlicht: (2026)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Evolving Deception: When Agents Evolve, Deception Wins
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions
von: Wang, Shenao, et al.
Veröffentlicht: (2026)
von: Wang, Shenao, et al.
Veröffentlicht: (2026)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
von: Jeong, Joonhyun, et al.
Veröffentlicht: (2025)
von: Jeong, Joonhyun, et al.
Veröffentlicht: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
von: Sun, Xiongtao, et al.
Veröffentlicht: (2025)
von: Sun, Xiongtao, et al.
Veröffentlicht: (2025)
DLP: towards active defense against backdoor attacks with decoupled learning process
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
NBA: defensive distillation for backdoor removal via neural behavior alignment
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
von: Yan, Song, et al.
Veröffentlicht: (2025)
von: Yan, Song, et al.
Veröffentlicht: (2025)
SoK: Understanding Vulnerabilities in the Large Language Model Supply Chain
von: Wang, Shenao, et al.
Veröffentlicht: (2025)
von: Wang, Shenao, et al.
Veröffentlicht: (2025)
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything
von: Zou, Xiaotian, et al.
Veröffentlicht: (2024)
von: Zou, Xiaotian, et al.
Veröffentlicht: (2024)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks
von: Zou, Quanchen, et al.
Veröffentlicht: (2026)
von: Zou, Quanchen, et al.
Veröffentlicht: (2026)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
von: Xiang, Shiyu, et al.
Veröffentlicht: (2025)
von: Xiang, Shiyu, et al.
Veröffentlicht: (2025)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
von: He, Zeqing, et al.
Veröffentlicht: (2024)
von: He, Zeqing, et al.
Veröffentlicht: (2024)
Multimodal Pragmatic Jailbreak on Text-to-image Models
von: Liu, Tong, et al.
Veröffentlicht: (2024)
von: Liu, Tong, et al.
Veröffentlicht: (2024)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
Mitigating Jailbreaks with Intent-Aware LLMs
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026) -
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
von: Chen, Moyang, et al.
Veröffentlicht: (2026) -
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
von: Zou, Quanchen, et al.
Veröffentlicht: (2025) -
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
von: Zou, Quanchen, et al.
Veröffentlicht: (2026) -
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
von: Zhang, Deyue, et al.
Veröffentlicht: (2025)