Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhexin, Yang, Junxiao, Ke, Pei, Mi, Fei, Wang, Hongning, Huang, Minlie
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910482538954752
author Zhang, Zhexin
Yang, Junxiao
Ke, Pei
Mi, Fei
Wang, Hongning
Huang, Minlie
author_facet Zhang, Zhexin
Yang, Junxiao
Ke, Pei
Mi, Fei
Wang, Hongning
Huang, Minlie
contents While significant attention has been dedicated to exploiting weaknesses in LLMs through jailbreaking attacks, there remains a paucity of effort in defending against these attacks. We point out a pivotal factor contributing to the success of jailbreaks: the intrinsic conflict between the goals of being helpful and ensuring safety. Accordingly, we propose to integrate goal prioritization at both training and inference stages to counteract. Implementing goal prioritization during inference substantially diminishes the Attack Success Rate (ASR) of jailbreaking from 66.4% to 3.6% for ChatGPT. And integrating goal prioritization into model training reduces the ASR from 71.0% to 6.6% for Llama2-13B. Remarkably, even in scenarios where no jailbreaking samples are included during training, our approach slashes the ASR by half. Additionally, our findings reveal that while stronger LLMs face greater safety risks, they also possess a greater capacity to be steered towards defending against such attacks, both because of their stronger ability in instruction following. Our work thus contributes to the comprehension of jailbreaking attacks and defenses, and sheds light on the relationship between LLMs' capability and safety. Our code is available at \url{https://github.com/thu-coai/JailbreakDefense_GoalPriority}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09096
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
Zhang, Zhexin
Yang, Junxiao
Ke, Pei
Mi, Fei
Wang, Hongning
Huang, Minlie
Computation and Language
While significant attention has been dedicated to exploiting weaknesses in LLMs through jailbreaking attacks, there remains a paucity of effort in defending against these attacks. We point out a pivotal factor contributing to the success of jailbreaks: the intrinsic conflict between the goals of being helpful and ensuring safety. Accordingly, we propose to integrate goal prioritization at both training and inference stages to counteract. Implementing goal prioritization during inference substantially diminishes the Attack Success Rate (ASR) of jailbreaking from 66.4% to 3.6% for ChatGPT. And integrating goal prioritization into model training reduces the ASR from 71.0% to 6.6% for Llama2-13B. Remarkably, even in scenarios where no jailbreaking samples are included during training, our approach slashes the ASR by half. Additionally, our findings reveal that while stronger LLMs face greater safety risks, they also possess a greater capacity to be steered towards defending against such attacks, both because of their stronger ability in instruction following. Our work thus contributes to the comprehension of jailbreaking attacks and defenses, and sheds light on the relationship between LLMs' capability and safety. Our code is available at \url{https://github.com/thu-coai/JailbreakDefense_GoalPriority}.
title Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
topic Computation and Language
url https://arxiv.org/abs/2311.09096