CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yuetai, Xu, Zhangchen, Jiang, Fengqing, Niu, Luyao, Sahabandu, Dinuka, Ramasubramanian, Bhaskar, Poovendran, Radha
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913761469661184
author Li, Yuetai
Xu, Zhangchen
Jiang, Fengqing
Niu, Luyao
Sahabandu, Dinuka
Ramasubramanian, Bhaskar
Poovendran, Radha
author_facet Li, Yuetai
Xu, Zhangchen
Jiang, Fengqing
Niu, Luyao
Sahabandu, Dinuka
Ramasubramanian, Bhaskar
Poovendran, Radha
contents The remarkable performance of large language models (LLMs) in generation tasks has enabled practitioners to leverage publicly available models to power custom applications, such as chatbots and virtual assistants. However, the data used to train or fine-tune these LLMs is often undisclosed, allowing an attacker to compromise the data and inject backdoors into the models. In this paper, we develop a novel inference time defense, named CLEANGEN, to mitigate backdoor attacks for generation tasks in LLMs. CLEANGEN is a lightweight and effective decoding strategy that is compatible with the state-of-the-art (SOTA) LLMs. Our insight behind CLEANGEN is that compared to other LLMs, backdoored LLMs assign significantly higher probabilities to tokens representing the attacker-desired contents. These discrepancies in token probabilities enable CLEANGEN to identify suspicious tokens favored by the attacker and replace them with tokens generated by another LLM that is not compromised by the same attacker, thereby avoiding generation of attacker-desired content. We evaluate CLEANGEN against five SOTA backdoor attacks. Our results show that CLEANGEN achieves lower attack success rates (ASR) compared to five SOTA baseline defenses for all five backdoor attacks. Moreover, LLMs deploying CLEANGEN maintain helpfulness in their responses when serving benign user queries with minimal added computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12257
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
Li, Yuetai
Xu, Zhangchen
Jiang, Fengqing
Niu, Luyao
Sahabandu, Dinuka
Ramasubramanian, Bhaskar
Poovendran, Radha
Artificial Intelligence
Cryptography and Security
The remarkable performance of large language models (LLMs) in generation tasks has enabled practitioners to leverage publicly available models to power custom applications, such as chatbots and virtual assistants. However, the data used to train or fine-tune these LLMs is often undisclosed, allowing an attacker to compromise the data and inject backdoors into the models. In this paper, we develop a novel inference time defense, named CLEANGEN, to mitigate backdoor attacks for generation tasks in LLMs. CLEANGEN is a lightweight and effective decoding strategy that is compatible with the state-of-the-art (SOTA) LLMs. Our insight behind CLEANGEN is that compared to other LLMs, backdoored LLMs assign significantly higher probabilities to tokens representing the attacker-desired contents. These discrepancies in token probabilities enable CLEANGEN to identify suspicious tokens favored by the attacker and replace them with tokens generated by another LLM that is not compromised by the same attacker, thereby avoiding generation of attacker-desired content. We evaluate CLEANGEN against five SOTA backdoor attacks. Our results show that CLEANGEN achieves lower attack success rates (ASR) compared to five SOTA baseline defenses for all five backdoor attacks. Moreover, LLMs deploying CLEANGEN maintain helpfulness in their responses when serving benign user queries with minimal added computational overhead.
title CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2406.12257