Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Ying, Zonghao, Liu, Aishan, Liu, Xianglong, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
by: Chen, Moyang, et al.
Published: (2026)
by: Chen, Moyang, et al.
Published: (2026)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
by: Ye, Mang, et al.
Published: (2025)
by: Ye, Mang, et al.
Published: (2025)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
by: Zou, Quanchen, et al.
Published: (2025)
by: Zou, Quanchen, et al.
Published: (2025)
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
by: Xu, Wenzhuo, et al.
Published: (2025)
by: Xu, Wenzhuo, et al.
Published: (2025)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
by: Gao, Sensen, et al.
Published: (2024)
by: Gao, Sensen, et al.
Published: (2024)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
by: Jing, Zonglei, et al.
Published: (2025)
by: Jing, Zonglei, et al.
Published: (2025)
Metaphor-based Jailbreak Attacks on Text-to-Image Models
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
by: Jia, Xiaojun, et al.
Published: (2025)
by: Jia, Xiaojun, et al.
Published: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
Responsible Diffusion: A Comprehensive Survey on Safety, Ethics, and Trust in Diffusion Models
by: Wei, Kang, et al.
Published: (2025)
by: Wei, Kang, et al.
Published: (2025)
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural Networks
by: Yu, Yi, et al.
Published: (2026)
by: Yu, Yi, et al.
Published: (2026)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
by: Liang, Shuang, et al.
Published: (2025)
by: Liang, Shuang, et al.
Published: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Intriguing Properties of Diffusion Models: An Empirical Study of the Natural Attack Capability in Text-to-Image Generative Models
by: Sato, Takami, et al.
Published: (2023)
by: Sato, Takami, et al.
Published: (2023)
Antelope: Potent and Concealed Jailbreak Attack Strategy
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
by: Xu, Shuhan, et al.
Published: (2026)
by: Xu, Shuhan, et al.
Published: (2026)
Anomaly Unveiled: Securing Image Classification against Adversarial Patch Attacks
by: Chattopadhyay, Nandish, et al.
Published: (2024)
by: Chattopadhyay, Nandish, et al.
Published: (2024)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
by: Liu, Jiangtao, et al.
Published: (2025)
by: Liu, Jiangtao, et al.
Published: (2025)
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models
by: Yang, Ying, et al.
Published: (2025)
by: Yang, Ying, et al.
Published: (2025)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
by: Zhong, Zhiyuan, et al.
Published: (2025)
by: Zhong, Zhiyuan, et al.
Published: (2025)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning
by: Komaty, Alain, et al.
Published: (2025)
by: Komaty, Alain, et al.
Published: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
by: Liu, Yilian, et al.
Published: (2026)
by: Liu, Yilian, et al.
Published: (2026)
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
How Generalizable are Deepfake Image Detectors? An Empirical Study
by: Li, Boquan, et al.
Published: (2023)
by: Li, Boquan, et al.
Published: (2023)
Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection
by: Cheng, Zhiyuan, et al.
Published: (2023)
by: Cheng, Zhiyuan, et al.
Published: (2023)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
by: Qiu, Yixiang, et al.
Published: (2025)
by: Qiu, Yixiang, et al.
Published: (2025)
Jailbreaking Attack against Multimodal Large Language Model
by: Niu, Zhenxing, et al.
Published: (2024)
by: Niu, Zhenxing, et al.
Published: (2024)
Similar Items
-
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
by: Ying, Zonghao, et al.
Published: (2024) -
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
by: Ying, Zonghao, et al.
Published: (2025) -
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
by: Chen, Moyang, et al.
Published: (2026) -
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
by: Ying, Zonghao, et al.
Published: (2025) -
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
by: Guo, Yangyang, et al.
Published: (2024)