Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liang, Zi, Ye, Qingqing, Liu, Xuan, Wang, Yanyun, Xu, Jianliang, Hu, Haibo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914110968430592
author Liang, Zi
Ye, Qingqing
Liu, Xuan
Wang, Yanyun
Xu, Jianliang
Hu, Haibo
author_facet Liang, Zi
Ye, Qingqing
Liu, Xuan
Wang, Yanyun
Xu, Jianliang
Hu, Haibo
contents Synthetic data refers to artificial samples generated by models. While it has been validated to significantly enhance the performance of large language models (LLMs) during training and has been widely adopted in LLM development, potential security risks it may introduce remain uninvestigated. This paper systematically evaluates the resilience of synthetic-data-integrated training paradigm for LLMs against mainstream poisoning and backdoor attacks. We reveal that such a paradigm exhibits strong resistance to existing attacks, primarily thanks to the different distribution patterns between poisoning data and queries used to generate synthetic samples. To enhance the effectiveness of these attacks and further investigate the security risks introduced by synthetic data, we introduce a novel and universal attack framework, namely, Virus Infection Attack (VIA), which enables the propagation of current attacks through synthetic data even under purely clean queries. Inspired by the principles of virus design in cybersecurity, VIA conceals the poisoning payload within a protective "shell" and strategically searches for optimal hijacking points in benign samples to maximize the likelihood of generating malicious content. Extensive experiments on both data poisoning and backdoor attacks show that VIA significantly increases the presence of poisoning content in synthetic data and correspondingly raises the attack success rate (ASR) on downstream models to levels comparable to those observed in the poisoned upstream models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23041
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
Liang, Zi
Ye, Qingqing
Liu, Xuan
Wang, Yanyun
Xu, Jianliang
Hu, Haibo
Cryptography and Security
Artificial Intelligence
Computation and Language
Synthetic data refers to artificial samples generated by models. While it has been validated to significantly enhance the performance of large language models (LLMs) during training and has been widely adopted in LLM development, potential security risks it may introduce remain uninvestigated. This paper systematically evaluates the resilience of synthetic-data-integrated training paradigm for LLMs against mainstream poisoning and backdoor attacks. We reveal that such a paradigm exhibits strong resistance to existing attacks, primarily thanks to the different distribution patterns between poisoning data and queries used to generate synthetic samples. To enhance the effectiveness of these attacks and further investigate the security risks introduced by synthetic data, we introduce a novel and universal attack framework, namely, Virus Infection Attack (VIA), which enables the propagation of current attacks through synthetic data even under purely clean queries. Inspired by the principles of virus design in cybersecurity, VIA conceals the poisoning payload within a protective "shell" and strategically searches for optimal hijacking points in benign samples to maximize the likelihood of generating malicious content. Extensive experiments on both data poisoning and backdoor attacks show that VIA significantly increases the presence of poisoning content in synthetic data and correspondingly raises the attack success rate (ASR) on downstream models to levels comparable to those observed in the poisoned upstream models.
title Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.23041