NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Chuhan, Zhang, Ye, Shi, Bowen, Gan, Yuyou, Du, Tianyu, Ji, Shouling, Deng, Dazhan, Wu, Yingcai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908518998605824
author Zhang, Chuhan
Zhang, Ye
Shi, Bowen
Gan, Yuyou
Du, Tianyu
Ji, Shouling
Deng, Dazhan
Wu, Yingcai
author_facet Zhang, Chuhan
Zhang, Ye
Shi, Bowen
Gan, Yuyou
Du, Tianyu
Ji, Shouling
Deng, Dazhan
Wu, Yingcai
contents In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanisms with adversarial prompts, has placed increasing pressure on the security defenses of LLMs. Strengthening resistance to jailbreak attacks requires an in-depth understanding of the security mechanisms and vulnerabilities of LLMs. However, the vast number of parameters and complex structure of LLMs make analyzing security weaknesses from an internal perspective a challenging task. This paper presents NeuroBreak, a top-down jailbreak analysis system designed to analyze neuron-level safety mechanisms and mitigate vulnerabilities. We carefully design system requirements through collaboration with three experts in the field of AI security. The system provides a comprehensive analysis of various jailbreak attack methods. By incorporating layer-wise representation probing analysis, NeuroBreak offers a novel perspective on the model's decision-making process throughout its generation steps. Furthermore, the system supports the analysis of critical neurons from both semantic and functional perspectives, facilitating a deeper exploration of security mechanisms. We conduct quantitative evaluations and case studies to verify the effectiveness of our system, offering mechanistic insights for developing next-generation defense strategies against evolving jailbreak attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03985
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
Zhang, Chuhan
Zhang, Ye
Shi, Bowen
Gan, Yuyou
Du, Tianyu
Ji, Shouling
Deng, Dazhan
Wu, Yingcai
Cryptography and Security
Artificial Intelligence
In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanisms with adversarial prompts, has placed increasing pressure on the security defenses of LLMs. Strengthening resistance to jailbreak attacks requires an in-depth understanding of the security mechanisms and vulnerabilities of LLMs. However, the vast number of parameters and complex structure of LLMs make analyzing security weaknesses from an internal perspective a challenging task. This paper presents NeuroBreak, a top-down jailbreak analysis system designed to analyze neuron-level safety mechanisms and mitigate vulnerabilities. We carefully design system requirements through collaboration with three experts in the field of AI security. The system provides a comprehensive analysis of various jailbreak attack methods. By incorporating layer-wise representation probing analysis, NeuroBreak offers a novel perspective on the model's decision-making process throughout its generation steps. Furthermore, the system supports the analysis of critical neurons from both semantic and functional perspectives, facilitating a deeper exploration of security mechanisms. We conduct quantitative evaluations and case studies to verify the effectiveness of our system, offering mechanistic insights for developing next-generation defense strategies against evolving jailbreak attacks.
title NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2509.03985