Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Xiaomeng, Chen, Pin-Yu, Ho, Tsung-Yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
por: Hu, Xiaomeng, et al.
Publicado: (2024)
por: Hu, Xiaomeng, et al.
Publicado: (2024)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
por: Ni, Ziyi, et al.
Publicado: (2025)
por: Ni, Ziyi, et al.
Publicado: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
por: Long, Zhuohang, et al.
Publicado: (2025)
por: Long, Zhuohang, et al.
Publicado: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
por: Pu, Rui, et al.
Publicado: (2024)
por: Pu, Rui, et al.
Publicado: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
por: Xu, Zhao, et al.
Publicado: (2024)
por: Xu, Zhao, et al.
Publicado: (2024)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
por: Ouyang, Yang, et al.
Publicado: (2025)
por: Ouyang, Yang, et al.
Publicado: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
por: Yi, Sibo, et al.
Publicado: (2024)
por: Yi, Sibo, et al.
Publicado: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)
por: Liu, Fan, et al.
Publicado: (2024)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
por: Akbar-Tajari, Mohammad, et al.
Publicado: (2025)
por: Akbar-Tajari, Mohammad, et al.
Publicado: (2025)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
por: Lv, Lijia, et al.
Publicado: (2024)
por: Lv, Lijia, et al.
Publicado: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
por: Chen, Taiye, et al.
Publicado: (2025)
por: Chen, Taiye, et al.
Publicado: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
por: Gohil, Vasudev
Publicado: (2025)
por: Gohil, Vasudev
Publicado: (2025)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
por: Kirch, Nathalie, et al.
Publicado: (2024)
por: Kirch, Nathalie, et al.
Publicado: (2024)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
por: Hu, Xiaomeng, et al.
Publicado: (2024)
por: Hu, Xiaomeng, et al.
Publicado: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
por: Chu, Junjie, et al.
Publicado: (2024)
por: Chu, Junjie, et al.
Publicado: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
por: Tu, Shangqing, et al.
Publicado: (2024)
por: Tu, Shangqing, et al.
Publicado: (2024)
Activation-Guided Local Editing for Jailbreaking Attacks
por: Wang, Jiecong, et al.
Publicado: (2025)
por: Wang, Jiecong, et al.
Publicado: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
por: Xiao, Zeguan, et al.
Publicado: (2024)
por: Xiao, Zeguan, et al.
Publicado: (2024)
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
por: Xu, Naen, et al.
Publicado: (2026)
por: Xu, Naen, et al.
Publicado: (2026)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
por: Yan, Dong, et al.
Publicado: (2026)
por: Yan, Dong, et al.
Publicado: (2026)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
por: Yu, Yongcan, et al.
Publicado: (2025)
por: Yu, Yongcan, et al.
Publicado: (2025)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
por: Xiong, Chen, et al.
Publicado: (2026)
por: Xiong, Chen, et al.
Publicado: (2026)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
por: Jiang, Weisen, et al.
Publicado: (2025)
por: Jiang, Weisen, et al.
Publicado: (2025)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
por: Yu, Miao, et al.
Publicado: (2024)
por: Yu, Miao, et al.
Publicado: (2024)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
por: Zhao, Yi, et al.
Publicado: (2025)
por: Zhao, Yi, et al.
Publicado: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
por: Mehrotra, Anay, et al.
Publicado: (2023)
por: Mehrotra, Anay, et al.
Publicado: (2023)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
por: Shang, Zhengchun, et al.
Publicado: (2025)
por: Shang, Zhengchun, et al.
Publicado: (2025)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
por: Xu, Zihao, et al.
Publicado: (2024)
por: Xu, Zihao, et al.
Publicado: (2024)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
por: Yi, Sibo, et al.
Publicado: (2025)
por: Yi, Sibo, et al.
Publicado: (2025)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
por: Yang, Yan, et al.
Publicado: (2024)
por: Yang, Yan, et al.
Publicado: (2024)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025)
por: Mu, Junjie, et al.
Publicado: (2025)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
por: Zhou, Guanghao, et al.
Publicado: (2025)
por: Zhou, Guanghao, et al.
Publicado: (2025)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
por: Zhao, Jiawei, et al.
Publicado: (2024)
por: Zhao, Jiawei, et al.
Publicado: (2024)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
por: Kabir, Md Rysul, et al.
Publicado: (2026)
por: Kabir, Md Rysul, et al.
Publicado: (2026)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
por: Gibbs, Tom, et al.
Publicado: (2024)
por: Gibbs, Tom, et al.
Publicado: (2024)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
por: Tong, Haibo, et al.
Publicado: (2025)
por: Tong, Haibo, et al.
Publicado: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
por: Li, Xirui, et al.
Publicado: (2024)
por: Li, Xirui, et al.
Publicado: (2024)
May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
por: Pandya, Nishit V., et al.
Publicado: (2025)
por: Pandya, Nishit V., et al.
Publicado: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
por: Liu, Xiaoqun, et al.
Publicado: (2024)
por: Liu, Xiaoqun, et al.
Publicado: (2024)
Ejemplares similares
-
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
por: Hu, Xiaomeng, et al.
Publicado: (2024) -
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
por: Ni, Ziyi, et al.
Publicado: (2025) -
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
por: Long, Zhuohang, et al.
Publicado: (2025) -
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
por: Pu, Rui, et al.
Publicado: (2024) -
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)