Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913021782130688 |
|---|---|
| author | Orgad, Hadas Wei, Boyi Zheng, Kaden Wattenberg, Martin Henderson, Peter Goldfarb-Tarrant, Seraphina Belinkov, Yonatan |
| author_facet | Orgad, Hadas Wei, Boyi Zheng, Kaden Wattenberg, Martin Henderson, Peter Goldfarb-Tarrant, Seraphina Belinkov, Yonatan |
| contents | Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly. Whether this brittleness reflects a fundamental lack of coherent internal organization for harmfulness remains unclear. Here we use targeted weight pruning as a causal intervention to probe the internal organization of harmfulness in LLMs. We find that harmful content generation depends on a compact set of weights that are general across harm types and distinct from benign capabilities. Aligned models exhibit a greater compression of harm generation weights than unaligned counterparts, indicating that alignment reshapes harmful representations internally--despite the brittleness of safety guardrails at the surface level. This compression explains emergent misalignment: if weights of harmful capabilities are compressed, fine-tuning that engages these weights in one domain can trigger broad misalignment. Consistent with this, pruning harm generation weights in a narrow domain substantially reduces emergent misalignment. Notably, LLMs harmful generation capability is dissociated from how they recognize and explain such content. Together, these results reveal a coherent internal structure for harmfulness in LLMs that may serve as a foundation for more principled approaches to safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_09544 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism Orgad, Hadas Wei, Boyi Zheng, Kaden Wattenberg, Martin Henderson, Peter Goldfarb-Tarrant, Seraphina Belinkov, Yonatan Computation and Language Artificial Intelligence Machine Learning I.2.7 Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly. Whether this brittleness reflects a fundamental lack of coherent internal organization for harmfulness remains unclear. Here we use targeted weight pruning as a causal intervention to probe the internal organization of harmfulness in LLMs. We find that harmful content generation depends on a compact set of weights that are general across harm types and distinct from benign capabilities. Aligned models exhibit a greater compression of harm generation weights than unaligned counterparts, indicating that alignment reshapes harmful representations internally--despite the brittleness of safety guardrails at the surface level. This compression explains emergent misalignment: if weights of harmful capabilities are compressed, fine-tuning that engages these weights in one domain can trigger broad misalignment. Consistent with this, pruning harm generation weights in a narrow domain substantially reduces emergent misalignment. Notably, LLMs harmful generation capability is dissociated from how they recognize and explain such content. Together, these results reveal a coherent internal structure for harmfulness in LLMs that may serve as a foundation for more principled approaches to safety. |
| title | Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism |
| topic | Computation and Language Artificial Intelligence Machine Learning I.2.7 |
| url | https://arxiv.org/abs/2604.09544 |