Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tianlong, Wang, Zhenghua, Liu, Wenhao, Wu, Muling, Dou, Shihan, Lv, Changze, Wang, Xiaohua, Zheng, Xiaoqing, Huang, Xuanjing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916623278931968
author Li, Tianlong
Wang, Zhenghua
Liu, Wenhao
Wu, Muling
Dou, Shihan
Lv, Changze
Wang, Xiaohua
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Li, Tianlong
Wang, Zhenghua
Liu, Wenhao
Wu, Muling
Dou, Shihan
Lv, Changze
Wang, Xiaohua
Zheng, Xiaoqing
Huang, Xuanjing
contents The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06824
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Li, Tianlong
Wang, Zhenghua
Liu, Wenhao
Wu, Muling
Dou, Shihan
Lv, Changze
Wang, Xiaohua
Zheng, Xiaoqing
Huang, Xuanjing
Computation and Language
Artificial Intelligence
The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community.
title Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.06824