Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Shuo, Han, Zhen, He, Bailan, Ding, Zifeng, Yu, Wenqian, Torr, Philip, Tresp, Volker, Gu, Jindong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
Deep Research Brings Deeper Harm
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
von: Wang, Zefeng, et al.
Veröffentlicht: (2024)
von: Wang, Zefeng, et al.
Veröffentlicht: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
Multimodal Pragmatic Jailbreak on Text-to-image Models
von: Liu, Tong, et al.
Veröffentlicht: (2024)
von: Liu, Tong, et al.
Veröffentlicht: (2024)
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
Reimagining Safety Alignment with An Image
von: Xia, Yifan, et al.
Veröffentlicht: (2025)
von: Xia, Yifan, et al.
Veröffentlicht: (2025)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Low-Resource Languages Jailbreak GPT-4
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
$PC^2$: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models
von: Choi, Wonwoo, et al.
Veröffentlicht: (2026)
von: Choi, Wonwoo, et al.
Veröffentlicht: (2026)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
von: Mukherjee, Kunal, et al.
Veröffentlicht: (2026)
von: Mukherjee, Kunal, et al.
Veröffentlicht: (2026)
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
von: Ramesh, Govind, et al.
Veröffentlicht: (2024)
von: Ramesh, Govind, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
BadGPT-4o: stripping safety finetuning from GPT models
von: Krupkina, Ekaterina, et al.
Veröffentlicht: (2024)
von: Krupkina, Ekaterina, et al.
Veröffentlicht: (2024)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
von: Xiang, Shiyu, et al.
Veröffentlicht: (2025)
von: Xiang, Shiyu, et al.
Veröffentlicht: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
von: Gao, Sensen, et al.
Veröffentlicht: (2024)
von: Gao, Sensen, et al.
Veröffentlicht: (2024)
EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection
von: Liu, Shigang, et al.
Veröffentlicht: (2024)
von: Liu, Shigang, et al.
Veröffentlicht: (2024)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
von: Liu, Xiaoqun, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqun, et al.
Veröffentlicht: (2024)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
von: Horal, Artur, et al.
Veröffentlicht: (2025)
von: Horal, Artur, et al.
Veröffentlicht: (2025)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
von: Yang, Fan
Veröffentlicht: (2025)
von: Yang, Fan
Veröffentlicht: (2025)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
von: Piet, Julien, et al.
Veröffentlicht: (2025)
von: Piet, Julien, et al.
Veröffentlicht: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
von: Pathade, Chetan
Veröffentlicht: (2025)
von: Pathade, Chetan
Veröffentlicht: (2025)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
von: Lu, Lin, et al.
Veröffentlicht: (2024)
von: Lu, Lin, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
von: Chen, Shuo, et al.
Veröffentlicht: (2025) -
Deep Research Brings Deeper Harm
von: Chen, Shuo, et al.
Veröffentlicht: (2025) -
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024) -
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
von: Wang, Zefeng, et al.
Veröffentlicht: (2024) -
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)