Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Chongwen, Ke, Yutong, Huang, Kaizhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Defending against Jailbreak through Early Exit Generation of Large Language Models
by: Zhao, Chongwen, et al.
Published: (2024)
by: Zhao, Chongwen, et al.
Published: (2024)
SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
by: Zhao, Qinjian, et al.
Published: (2025)
by: Zhao, Qinjian, et al.
Published: (2025)
Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026)
by: Shen, Sicheng, et al.
Published: (2026)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Partial Differential Equations is All You Need for Generating Neural Architectures -- A Theory for Physical Artificial Intelligence Systems
by: Guo, Ping, et al.
Published: (2021)
by: Guo, Ping, et al.
Published: (2021)
Jailbreaking LLM-Controlled Robots
by: Robey, Alexander, et al.
Published: (2024)
by: Robey, Alexander, et al.
Published: (2024)
LLM-Powered Explanations: Unraveling Recommendations Through Subgraph Reasoning
by: Shi, Guangsi, et al.
Published: (2024)
by: Shi, Guangsi, et al.
Published: (2024)
Learning from Risk: LLM-Guided Generation of Safety-Critical Scenarios with Prior Knowledge
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
by: Li, Miles Q., et al.
Published: (2026)
by: Li, Miles Q., et al.
Published: (2026)
A generalizable framework for low-rank tensor completion with numerical priors
by: Yuan, Shiran, et al.
Published: (2023)
by: Yuan, Shiran, et al.
Published: (2023)
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
by: Hu, Yijie, et al.
Published: (2025)
by: Hu, Yijie, et al.
Published: (2025)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
by: Koo, Hamin, et al.
Published: (2025)
by: Koo, Hamin, et al.
Published: (2025)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
by: Wang, Fengxiang, et al.
Published: (2024)
by: Wang, Fengxiang, et al.
Published: (2024)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025)
by: Xin, Yuan, et al.
Published: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
The Echo Chamber Multi-Turn LLM Jailbreak
by: Alobaid, Ahmad, et al.
Published: (2026)
by: Alobaid, Ahmad, et al.
Published: (2026)
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
by: Zhao, Yinzhi, et al.
Published: (2026)
by: Zhao, Yinzhi, et al.
Published: (2026)
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025)
by: Kritz, Jeremy, et al.
Published: (2025)
Rethinking Multi-domain Generalization with A General Learning Objective
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
by: Zhang, Yingjie, et al.
Published: (2025)
by: Zhang, Yingjie, et al.
Published: (2025)
NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
by: Zhu, Hanyu, et al.
Published: (2025)
by: Zhu, Hanyu, et al.
Published: (2025)
IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons
by: Shi, Dan, et al.
Published: (2024)
by: Shi, Dan, et al.
Published: (2024)
MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
by: Zhang, Jianyi, et al.
Published: (2025)
by: Zhang, Jianyi, et al.
Published: (2025)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
by: Wang, Ziwei, et al.
Published: (2026)
by: Wang, Ziwei, et al.
Published: (2026)
SafeDream: Safety World Model for Proactive Early Jailbreak Detection
by: Yan, Bo, et al.
Published: (2026)
by: Yan, Bo, et al.
Published: (2026)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
LLM Jailbreak Detection for (Almost) Free!
by: Chen, Guorui, et al.
Published: (2025)
by: Chen, Guorui, et al.
Published: (2025)
Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language Models
by: Chen, Yuheng, et al.
Published: (2024)
by: Chen, Yuheng, et al.
Published: (2024)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
by: Zahran, Noureldin, et al.
Published: (2025)
by: Zahran, Noureldin, et al.
Published: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Similar Items
-
Defending against Jailbreak through Early Exit Generation of Large Language Models
by: Zhao, Chongwen, et al.
Published: (2024) -
SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
by: Zhao, Qinjian, et al.
Published: (2025) -
Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
by: Peng, Jingyu, et al.
Published: (2025) -
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024) -
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026)