BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Maozhen, Zhao, Mengnan, Wang, Wei, Wang, Bo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914025373171712
author Zhang, Maozhen
Zhao, Mengnan
Wang, Wei
Wang, Bo
author_facet Zhang, Maozhen
Zhao, Mengnan
Wang, Wei
Wang, Bo
contents Prompt-based tuning has emerged as a lightweight alternative to full fine-tuning in large vision-language models, enabling efficient adaptation via learned contextual prompts. This paradigm has recently been extended to federated learning settings (e.g., PromptFL), where clients collaboratively train prompts under data privacy constraints. However, the security implications of prompt-based aggregation in federated multimodal learning remain largely unexplored, leaving a critical attack surface unaddressed. In this paper, we introduce \textbf{BadPromptFL}, the first backdoor attack targeting prompt-based federated learning in multimodal contrastive models. In BadPromptFL, compromised clients jointly optimize local backdoor triggers and prompt embeddings, injecting poisoned prompts into the global aggregation process. These prompts are then propagated to benign clients, enabling universal backdoor activation at inference without modifying model parameters. Leveraging the contextual learning behavior of CLIP-style architectures, BadPromptFL achieves high attack success rates (e.g., \(>90\%\)) with minimal visibility and limited client participation. Extensive experiments across multiple datasets and aggregation protocols validate the effectiveness, stealth, and generalizability of our attack, raising critical concerns about the robustness of prompt-based federated learning in real-world deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models
Zhang, Maozhen
Zhao, Mengnan
Wang, Wei
Wang, Bo
Machine Learning
Artificial Intelligence
Prompt-based tuning has emerged as a lightweight alternative to full fine-tuning in large vision-language models, enabling efficient adaptation via learned contextual prompts. This paradigm has recently been extended to federated learning settings (e.g., PromptFL), where clients collaboratively train prompts under data privacy constraints. However, the security implications of prompt-based aggregation in federated multimodal learning remain largely unexplored, leaving a critical attack surface unaddressed. In this paper, we introduce \textbf{BadPromptFL}, the first backdoor attack targeting prompt-based federated learning in multimodal contrastive models. In BadPromptFL, compromised clients jointly optimize local backdoor triggers and prompt embeddings, injecting poisoned prompts into the global aggregation process. These prompts are then propagated to benign clients, enabling universal backdoor activation at inference without modifying model parameters. Leveraging the contextual learning behavior of CLIP-style architectures, BadPromptFL achieves high attack success rates (e.g., \(>90\%\)) with minimal visibility and limited client participation. Extensive experiments across multiple datasets and aggregation protocols validate the effectiveness, stealth, and generalizability of our attack, raising critical concerns about the robustness of prompt-based federated learning in real-world deployments.
title BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.08040