Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yuanwei, Li, Xiang, Liu, Yixin, Zhou, Pan, Sun, Lichao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Low-Resource Languages Jailbreak GPT-4
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
New Wide-Net-Casting Jailbreak Attacks Risk Large Models
von: Xiang, Qiuchi, et al.
Veröffentlicht: (2026)
von: Xiang, Qiuchi, et al.
Veröffentlicht: (2026)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
von: Shang, Yingjia, et al.
Veröffentlicht: (2025)
von: Shang, Yingjia, et al.
Veröffentlicht: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
A General Black-box Adversarial Attack on Graph-based Fake News Detectors
von: Zhu, Peican, et al.
Veröffentlicht: (2024)
von: Zhu, Peican, et al.
Veröffentlicht: (2024)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models
von: Luo, Haozheng, et al.
Veröffentlicht: (2025)
von: Luo, Haozheng, et al.
Veröffentlicht: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
von: Yan, Yu, et al.
Veröffentlicht: (2026)
von: Yan, Yu, et al.
Veröffentlicht: (2026)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2024)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2024)
Jailbreaking with Universal Multi-Prompts
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
von: Hackett, William, et al.
Veröffentlicht: (2025)
von: Hackett, William, et al.
Veröffentlicht: (2025)
Untargeted Adversarial Attack on Knowledge Graph Embeddings
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2024)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2024)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Quantifying the Noise of Structural Perturbations on Graph Adversarial Attacks
von: Fang, Junyuan, et al.
Veröffentlicht: (2025)
von: Fang, Junyuan, et al.
Veröffentlicht: (2025)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
von: Zeng, Qirun, et al.
Veröffentlicht: (2025)
von: Zeng, Qirun, et al.
Veröffentlicht: (2025)
Adversarial Attacks on Reinforcement Learning-based Medical Questionnaire Systems: Input-level Perturbation Strategies and Medical Constraint Validation
von: Liu, Peizhuo
Veröffentlicht: (2025)
von: Liu, Peizhuo
Veröffentlicht: (2025)
Ähnliche Einträge
-
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025) -
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024) -
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024) -
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026) -
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)