Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yu, Liu, Yuanhao, Cao, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
by: Xie, Yuanbo, et al.
Published: (2025)
by: Xie, Yuanbo, et al.
Published: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
by: Li, Hongyi, et al.
Published: (2024)
by: Li, Hongyi, et al.
Published: (2024)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
by: Zhang, Jiawen, et al.
Published: (2025)
by: Zhang, Jiawen, et al.
Published: (2025)
The Privacy-Utility Trade-off in the Topics API
by: Alvim, Mário S., et al.
Published: (2024)
by: Alvim, Mário S., et al.
Published: (2024)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
by: Wang, Ziwei, et al.
Published: (2026)
by: Wang, Ziwei, et al.
Published: (2026)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
by: Cui, Tiehan, et al.
Published: (2025)
by: Cui, Tiehan, et al.
Published: (2025)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
A Unified Learn-to-Distort-Data Framework for Privacy-Utility Trade-off in Trustworthy Federated Learning
by: Zhang, Xiaojin, et al.
Published: (2024)
by: Zhang, Xiaojin, et al.
Published: (2024)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Clients Collaborate: Flexible Differentially Private Federated Learning with Guaranteed Improvement of Utility-Privacy Trade-off
by: Li, Yuecheng, et al.
Published: (2024)
by: Li, Yuecheng, et al.
Published: (2024)
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
by: Hammadia, Taha, et al.
Published: (2026)
by: Hammadia, Taha, et al.
Published: (2026)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
by: Cheng, Siyang, et al.
Published: (2025)
by: Cheng, Siyang, et al.
Published: (2025)
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
by: Collu, Matteo Gioele, et al.
Published: (2026)
by: Collu, Matteo Gioele, et al.
Published: (2026)
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
by: Ahn, Yelim, et al.
Published: (2025)
by: Ahn, Yelim, et al.
Published: (2025)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
by: Li, Qinfeng, et al.
Published: (2025)
by: Li, Qinfeng, et al.
Published: (2025)
Revisiting Privacy-Utility Trade-off for DP Training with Pre-existing Knowledge
by: Zheng, Yu, et al.
Published: (2024)
by: Zheng, Yu, et al.
Published: (2024)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
by: Tan, Yixin, et al.
Published: (2025)
by: Tan, Yixin, et al.
Published: (2025)
SafeDream: Safety World Model for Proactive Early Jailbreak Detection
by: Yan, Bo, et al.
Published: (2026)
by: Yan, Bo, et al.
Published: (2026)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
by: Wu, Daoyuan, et al.
Published: (2024)
by: Wu, Daoyuan, et al.
Published: (2024)
Similar Items
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024) -
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025) -
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
by: Xie, Yuanbo, et al.
Published: (2025) -
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026) -
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)