LoBAM: LoRA-Based Backdoor Attack on Model Merging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Ming, Zhang, Jingyang, Sun, Jingwei, Fang, Minghong, Li, Hai, Chen, Yiran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915313401987072
author Yin, Ming
Zhang, Jingyang
Sun, Jingwei
Fang, Minghong
Li, Hai
Chen, Yiran
author_facet Yin, Ming
Zhang, Jingyang
Sun, Jingwei
Fang, Minghong
Li, Hai
Chen, Yiran
contents Model merging is an emerging technique that integrates multiple models fine-tuned on different tasks to create a versatile model that excels in multiple domains. This scheme, in the meantime, may open up backdoor attack opportunities where one single malicious model can jeopardize the integrity of the merged model. Existing works try to demonstrate the risk of such attacks by assuming substantial computational resources, focusing on cases where the attacker can fully fine-tune the pre-trained model. Such an assumption, however, may not be feasible given the increasing size of machine learning models. In practice where resources are limited and the attacker can only employ techniques like Low-Rank Adaptation (LoRA) to produce the malicious model, it remains unclear whether the attack can still work and pose threats. In this work, we first identify that the attack efficacy is significantly diminished when using LoRA for fine-tuning. Then, we propose LoBAM, a method that yields high attack success rate with minimal training resources. The key idea of LoBAM is to amplify the malicious weights in an intelligent way that effectively enhances the attack efficacy. We demonstrate that our design can lead to improved attack success rate through extensive empirical experiments across various model merging scenarios. Moreover, we show that our method is highly stealthy and is difficult to detect and defend against.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16746
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LoBAM: LoRA-Based Backdoor Attack on Model Merging
Yin, Ming
Zhang, Jingyang
Sun, Jingwei
Fang, Minghong
Li, Hai
Chen, Yiran
Cryptography and Security
Artificial Intelligence
Machine Learning
Model merging is an emerging technique that integrates multiple models fine-tuned on different tasks to create a versatile model that excels in multiple domains. This scheme, in the meantime, may open up backdoor attack opportunities where one single malicious model can jeopardize the integrity of the merged model. Existing works try to demonstrate the risk of such attacks by assuming substantial computational resources, focusing on cases where the attacker can fully fine-tune the pre-trained model. Such an assumption, however, may not be feasible given the increasing size of machine learning models. In practice where resources are limited and the attacker can only employ techniques like Low-Rank Adaptation (LoRA) to produce the malicious model, it remains unclear whether the attack can still work and pose threats. In this work, we first identify that the attack efficacy is significantly diminished when using LoRA for fine-tuning. Then, we propose LoBAM, a method that yields high attack success rate with minimal training resources. The key idea of LoBAM is to amplify the malicious weights in an intelligent way that effectively enhances the attack efficacy. We demonstrate that our design can lead to improved attack success rate through extensive empirical experiments across various model merging scenarios. Moreover, we show that our method is highly stealthy and is difficult to detect and defend against.
title LoBAM: LoRA-Based Backdoor Attack on Model Merging
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.16746