Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Weixiang, Hu, Yulin, Li, Zhuojun, Deng, Yang, Guo, Jiahe, Sui, Xingyu, Zhao, Yanyan, Qin, Bing, Chua, Tat-Seng, Liu, Ting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913614413168640
author Zhao, Weixiang
Hu, Yulin
Li, Zhuojun
Deng, Yang
Guo, Jiahe
Sui, Xingyu
Zhao, Yanyan
Qin, Bing
Chua, Tat-Seng
Liu, Ting
author_facet Zhao, Weixiang
Hu, Yulin
Li, Zhuojun
Deng, Yang
Guo, Jiahe
Sui, Xingyu
Zhao, Yanyan
Qin, Bing
Chua, Tat-Seng
Liu, Ting
contents Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generate unsafe responses, exhibit over-safety by rejecting safe user inputs, and fail to preserve general utility after safety alignment. To this end, we propose a novel post safety alignment (PSA) method to address these inherent and emerging safety challenges, including safety enhancement, over-safety mitigation, and utility preservation. In specific, we introduce \textsc{SafePatching}, a novel framework for comprehensive PSA, where two distinct safety patches are developed on the harmful data to enhance safety and mitigate over-safety concerns, and then seamlessly integrated into the target LLM backbone without compromising its utility. Extensive experiments on four representative aligned LLMs, including LLaMA-2/3, Gemma and Mistral, show that \textsc{SafePatching} achieves a more comprehensive PSA than baseline methods, further optimizing the balance between being helpful and harmless in current aligned LLMs. Also, \textsc{SafePatching} demonstrates its superiority in continual PSA scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13820
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
Zhao, Weixiang
Hu, Yulin
Li, Zhuojun
Deng, Yang
Guo, Jiahe
Sui, Xingyu
Zhao, Yanyan
Qin, Bing
Chua, Tat-Seng
Liu, Ting
Computation and Language
Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generate unsafe responses, exhibit over-safety by rejecting safe user inputs, and fail to preserve general utility after safety alignment. To this end, we propose a novel post safety alignment (PSA) method to address these inherent and emerging safety challenges, including safety enhancement, over-safety mitigation, and utility preservation. In specific, we introduce \textsc{SafePatching}, a novel framework for comprehensive PSA, where two distinct safety patches are developed on the harmful data to enhance safety and mitigate over-safety concerns, and then seamlessly integrated into the target LLM backbone without compromising its utility. Extensive experiments on four representative aligned LLMs, including LLaMA-2/3, Gemma and Mistral, show that \textsc{SafePatching} achieves a more comprehensive PSA than baseline methods, further optimizing the balance between being helpful and harmless in current aligned LLMs. Also, \textsc{SafePatching} demonstrates its superiority in continual PSA scenarios.
title Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
topic Computation and Language
url https://arxiv.org/abs/2405.13820