MEGen: Generative Backdoor into Large Language Models via Model Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qiu, Jiyang, Ma, Xinbei, Zhang, Zhuosheng, Zhao, Hai, Li, Yun, Wang, Qianren
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909761578991616
author Qiu, Jiyang
Ma, Xinbei
Zhang, Zhuosheng
Zhao, Hai
Li, Yun
Wang, Qianren
author_facet Qiu, Jiyang
Ma, Xinbei
Zhang, Zhuosheng
Zhao, Hai
Li, Yun
Wang, Qianren
contents Large language models (LLMs) have exhibited remarkable versatility and adaptability, while their widespread adoption across various applications also raises critical safety concerns. This paper focuses on the impact of backdoored LLMs. Traditional backdoor injection methods are primarily limited to yes-or-no discriminative tasks, leading users to underestimate the potential risks of backdoored LLMs. Given the inherently generative nature of LLMs, this paper reveals that a generative backdoor injected into LLMs can expose the true safety risks in their applications. We propose an editing-based generative backdoor, named MEGen, aiming to expand the backdoor to generative tasks in a unified format of any text-to any text, leading to natural generations with a specific intention. Experiments show that MEGen achieves a high attack success rate by adjusting only a small set of local parameters with few-shot samples. Notably, we show that the backdoored model, when triggered, can freely output pre-set dangerous information while completing downstream tasks. Our work highlights that MEGen enables backdoors in LLMs to exhibit generative capabilities, causing potential safety risks by altering the generative style. The code is available at https://github.com/MonoQ-hub/MEGen.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10722
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MEGen: Generative Backdoor into Large Language Models via Model Editing
Qiu, Jiyang
Ma, Xinbei
Zhang, Zhuosheng
Zhao, Hai
Li, Yun
Wang, Qianren
Computation and Language
Artificial Intelligence
Large language models (LLMs) have exhibited remarkable versatility and adaptability, while their widespread adoption across various applications also raises critical safety concerns. This paper focuses on the impact of backdoored LLMs. Traditional backdoor injection methods are primarily limited to yes-or-no discriminative tasks, leading users to underestimate the potential risks of backdoored LLMs. Given the inherently generative nature of LLMs, this paper reveals that a generative backdoor injected into LLMs can expose the true safety risks in their applications. We propose an editing-based generative backdoor, named MEGen, aiming to expand the backdoor to generative tasks in a unified format of any text-to any text, leading to natural generations with a specific intention. Experiments show that MEGen achieves a high attack success rate by adjusting only a small set of local parameters with few-shot samples. Notably, we show that the backdoored model, when triggered, can freely output pre-set dangerous information while completing downstream tasks. Our work highlights that MEGen enables backdoors in LLMs to exhibit generative capabilities, causing potential safety risks by altering the generative style. The code is available at https://github.com/MonoQ-hub/MEGen.
title MEGen: Generative Backdoor into Large Language Models via Model Editing
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2408.10722