What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kirch, Nathalie, Weisser, Constantin, Field, Severin, Yannakoudakis, Helen, Casper, Stephen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
von: Pu, Rui, et al.
Veröffentlicht: (2024)
von: Pu, Rui, et al.
Veröffentlicht: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
von: Gohil, Vasudev
Veröffentlicht: (2025)
von: Gohil, Vasudev
Veröffentlicht: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
von: Ostermann, Simon, et al.
Veröffentlicht: (2024)
von: Ostermann, Simon, et al.
Veröffentlicht: (2024)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
von: Lv, Lijia, et al.
Veröffentlicht: (2024)
von: Lv, Lijia, et al.
Veröffentlicht: (2024)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
Activation-Guided Local Editing for Jailbreaking Attacks
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
von: Weiss, Roy, et al.
Veröffentlicht: (2024)
von: Weiss, Roy, et al.
Veröffentlicht: (2024)
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2024)
von: Yu, Miao, et al.
Veröffentlicht: (2024)
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
von: Song, Xurui, et al.
Veröffentlicht: (2025)
von: Song, Xurui, et al.
Veröffentlicht: (2025)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
von: Berezin, Sergey, et al.
Veröffentlicht: (2025)
von: Berezin, Sergey, et al.
Veröffentlicht: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
von: Ji, Wence, et al.
Veröffentlicht: (2025)
von: Ji, Wence, et al.
Veröffentlicht: (2025)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
von: Shen, Guangyu, et al.
Veröffentlicht: (2024)
von: Shen, Guangyu, et al.
Veröffentlicht: (2024)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Jailbreaking with Universal Multi-Prompts
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
von: Ouyang, Yang, et al.
Veröffentlicht: (2025)
von: Ouyang, Yang, et al.
Veröffentlicht: (2025)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
von: Gong, Yichen, et al.
Veröffentlicht: (2023)
von: Gong, Yichen, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024) -
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024) -
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
von: Li, Xirui, et al.
Veröffentlicht: (2024) -
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)