Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cao, Yuanpu, Cao, Bochuan, Chen, Jinghui
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910477112573952
author Cao, Yuanpu
Cao, Bochuan
Chen, Jinghui
author_facet Cao, Yuanpu
Cao, Bochuan
Chen, Jinghui
contents Recent developments in Large Language Models (LLMs) have manifested significant advancements. To facilitate safeguards against malicious exploitation, a body of research has concentrated on aligning LLMs with human preferences and inhibiting their generation of inappropriate content. Unfortunately, such alignments are often vulnerable: fine-tuning with a minimal amount of harmful data can easily unalign the target LLM. While being effective, such fine-tuning-based unalignment approaches also have their own limitations: (1) non-stealthiness, after fine-tuning, safety audits or red-teaming can easily expose the potential weaknesses of the unaligned models, thereby precluding their release/use. (2) non-persistence, the unaligned LLMs can be easily repaired through re-alignment, i.e., fine-tuning again with aligned data points. In this work, we show that it is possible to conduct stealthy and persistent unalignment on large language models via backdoor injections. We also provide a novel understanding on the relationship between the backdoor persistence and the activation pattern and further provide guidelines for potential trigger design. Through extensive experiments, we demonstrate that our proposed stealthy and persistent unalignment can successfully pass the safety evaluation while maintaining strong persistence against re-alignment defense.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00027
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
Cao, Yuanpu
Cao, Bochuan
Chen, Jinghui
Cryptography and Security
Artificial Intelligence
Computation and Language
Recent developments in Large Language Models (LLMs) have manifested significant advancements. To facilitate safeguards against malicious exploitation, a body of research has concentrated on aligning LLMs with human preferences and inhibiting their generation of inappropriate content. Unfortunately, such alignments are often vulnerable: fine-tuning with a minimal amount of harmful data can easily unalign the target LLM. While being effective, such fine-tuning-based unalignment approaches also have their own limitations: (1) non-stealthiness, after fine-tuning, safety audits or red-teaming can easily expose the potential weaknesses of the unaligned models, thereby precluding their release/use. (2) non-persistence, the unaligned LLMs can be easily repaired through re-alignment, i.e., fine-tuning again with aligned data points. In this work, we show that it is possible to conduct stealthy and persistent unalignment on large language models via backdoor injections. We also provide a novel understanding on the relationship between the backdoor persistence and the activation pattern and further provide guidelines for potential trigger design. Through extensive experiments, we demonstrate that our proposed stealthy and persistent unalignment can successfully pass the safety evaluation while maintaining strong persistence against re-alignment defense.
title Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2312.00027