Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Heegyu, Yuk, Sehyun, Cho, Hyunsouk
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917598833147904
author Kim, Heegyu
Yuk, Sehyun
Cho, Hyunsouk
author_facet Kim, Heegyu
Yuk, Sehyun
Cho, Hyunsouk
contents Caution: This paper includes offensive words that could potentially cause unpleasantness. Language models (LMs) are vulnerable to exploitation for adversarial misuse. Training LMs for safety alignment is extensive and makes it hard to respond to fast-developing attacks immediately, such as jailbreaks. We propose self-refine with formatting that achieves outstanding safety even in non-safety-aligned LMs and evaluate our method alongside several defense baselines, demonstrating that it is the safest training-free method against jailbreak attacks. Additionally, we proposed a formatting method that improves the efficiency of the self-refine process while reducing attack success rates in fewer iterations. We've also observed that non-safety-aligned LMs outperform safety-aligned LMs in safety tasks by giving more helpful and safe responses. In conclusion, our findings can achieve less safety risk with fewer computational costs, allowing non-safety LM to be easily utilized in real-world service.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15180
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
Kim, Heegyu
Yuk, Sehyun
Cho, Hyunsouk
Machine Learning
Computation and Language
Cryptography and Security
Caution: This paper includes offensive words that could potentially cause unpleasantness. Language models (LMs) are vulnerable to exploitation for adversarial misuse. Training LMs for safety alignment is extensive and makes it hard to respond to fast-developing attacks immediately, such as jailbreaks. We propose self-refine with formatting that achieves outstanding safety even in non-safety-aligned LMs and evaluate our method alongside several defense baselines, demonstrating that it is the safest training-free method against jailbreak attacks. Additionally, we proposed a formatting method that improves the efficiency of the self-refine process while reducing attack success rates in fewer iterations. We've also observed that non-safety-aligned LMs outperform safety-aligned LMs in safety tasks by giving more helpful and safe responses. In conclusion, our findings can achieve less safety risk with fewer computational costs, allowing non-safety LM to be easily utilized in real-world service.
title Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2402.15180