Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yuan, Zenghui, Xu, Yangming, Shi, Jiawen, Zhou, Pan, Sun, Lichao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910973821976576
author Yuan, Zenghui
Xu, Yangming
Shi, Jiawen
Zhou, Pan
Sun, Lichao
author_facet Yuan, Zenghui
Xu, Yangming
Shi, Jiawen
Zhou, Pan
Sun, Lichao
contents Model merging for Large Language Models (LLMs) directly fuses the parameters of different models finetuned on various tasks, creating a unified model for multi-domain tasks. However, due to potential vulnerabilities in models available on open-source platforms, model merging is susceptible to backdoor attacks. In this paper, we propose Merge Hijacking, the first backdoor attack targeting model merging in LLMs. The attacker constructs a malicious upload model and releases it. Once a victim user merges it with any other models, the resulting merged model inherits the backdoor while maintaining utility across tasks. Merge Hijacking defines two main objectives-effectiveness and utility-and achieves them through four steps. Extensive experiments demonstrate the effectiveness of our attack across different models, merging algorithms, and tasks. Additionally, we show that the attack remains effective even when merging real-world models. Moreover, our attack demonstrates robustness against two inference-time defenses (Paraphrasing and CLEANGEN) and one training-time defense (Fine-pruning).
format Preprint
id arxiv_https___arxiv_org_abs_2505_23561
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
Yuan, Zenghui
Xu, Yangming
Shi, Jiawen
Zhou, Pan
Sun, Lichao
Cryptography and Security
Model merging for Large Language Models (LLMs) directly fuses the parameters of different models finetuned on various tasks, creating a unified model for multi-domain tasks. However, due to potential vulnerabilities in models available on open-source platforms, model merging is susceptible to backdoor attacks. In this paper, we propose Merge Hijacking, the first backdoor attack targeting model merging in LLMs. The attacker constructs a malicious upload model and releases it. Once a victim user merges it with any other models, the resulting merged model inherits the backdoor while maintaining utility across tasks. Merge Hijacking defines two main objectives-effectiveness and utility-and achieves them through four steps. Extensive experiments demonstrate the effectiveness of our attack across different models, merging algorithms, and tasks. Additionally, we show that the attack remains effective even when merging real-world models. Moreover, our attack demonstrates robustness against two inference-time defenses (Paraphrasing and CLEANGEN) and one training-time defense (Fine-pruning).
title Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
topic Cryptography and Security
url https://arxiv.org/abs/2505.23561