Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916623278931968 |
|---|---|
| author | Li, Tianlong Wang, Zhenghua Liu, Wenhao Wu, Muling Dou, Shihan Lv, Changze Wang, Xiaohua Zheng, Xiaoqing Huang, Xuanjing |
| author_facet | Li, Tianlong Wang, Zhenghua Liu, Wenhao Wu, Muling Dou, Shihan Lv, Changze Wang, Xiaohua Zheng, Xiaoqing Huang, Xuanjing |
| contents | The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_06824 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective Li, Tianlong Wang, Zhenghua Liu, Wenhao Wu, Muling Dou, Shihan Lv, Changze Wang, Xiaohua Zheng, Xiaoqing Huang, Xuanjing Computation and Language Artificial Intelligence The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community. |
| title | Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2401.06824 |