FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Singha, Mainak, Roy, Subhankar, Mehrotra, Sarthak, Jha, Ankit, Abdar, Moloud, Banerjee, Biplab, Ricci, Elisa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915473849843712
author Singha, Mainak
Roy, Subhankar
Mehrotra, Sarthak
Jha, Ankit
Abdar, Moloud
Banerjee, Biplab
Ricci, Elisa
author_facet Singha, Mainak
Roy, Subhankar
Mehrotra, Sarthak
Jha, Ankit
Abdar, Moloud
Banerjee, Biplab
Ricci, Elisa
contents In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. However, textual prompt tuning suffers from overfitting to known concepts, limiting its generalizability to unseen concepts. To address this limitation, we propose Multimodal Visual Prompt Tuning (FedMVP) that conditions the prompts on multimodal contextual information - derived from the input image and textual attribute features of a class. At the core of FedMVP is a PromptFormer module that synergistically aligns textual and visual features through a cross-attention mechanism. The dynamically generated multimodal visual prompts are then input to the frozen vision encoder of CLIP, and trained with a combination of CLIP similarity loss and a consistency loss. Extensive evaluation on 20 datasets, spanning three generalization settings, demonstrates that FedMVP not only preserves performance on in-distribution classes and domains, but also displays higher generalizability to unseen classes and domains, surpassing state-of-the-art methods by a notable margin of +1.57% - 2.26%. Code is available at https://github.com/mainaksingha01/FedMVP.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
Singha, Mainak
Roy, Subhankar
Mehrotra, Sarthak
Jha, Ankit
Abdar, Moloud
Banerjee, Biplab
Ricci, Elisa
Computer Vision and Pattern Recognition
In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. However, textual prompt tuning suffers from overfitting to known concepts, limiting its generalizability to unseen concepts. To address this limitation, we propose Multimodal Visual Prompt Tuning (FedMVP) that conditions the prompts on multimodal contextual information - derived from the input image and textual attribute features of a class. At the core of FedMVP is a PromptFormer module that synergistically aligns textual and visual features through a cross-attention mechanism. The dynamically generated multimodal visual prompts are then input to the frozen vision encoder of CLIP, and trained with a combination of CLIP similarity loss and a consistency loss. Extensive evaluation on 20 datasets, spanning three generalization settings, demonstrates that FedMVP not only preserves performance on in-distribution classes and domains, but also displays higher generalizability to unseen classes and domains, surpassing state-of-the-art methods by a notable margin of +1.57% - 2.26%. Code is available at https://github.com/mainaksingha01/FedMVP.
title FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.20860