Open-Vocabulary Federated Learning with Multimodal Prototyping

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zeng, Huimin, Yue, Zhenrui, Wang, Dong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913294998044672
author Zeng, Huimin
Yue, Zhenrui
Wang, Dong
author_facet Zeng, Huimin
Yue, Zhenrui
Wang, Dong
contents Existing federated learning (FL) studies usually assume the training label space and test label space are identical. However, in real-world applications, this assumption is too ideal to be true. A new user could come up with queries that involve data from unseen classes, and such open-vocabulary queries would directly defect such FL systems. Therefore, in this work, we explicitly focus on the under-explored open-vocabulary challenge in FL. That is, for a new user, the global server shall understand her/his query that involves arbitrary unknown classes. To address this problem, we leverage the pre-trained vision-language models (VLMs). In particular, we present a novel adaptation framework tailored for VLMs in the context of FL, named as Federated Multimodal Prototyping (Fed-MP). Fed-MP adaptively aggregates the local model weights based on light-weight client residuals, and makes predictions based on a novel multimodal prototyping mechanism. Fed-MP exploits the knowledge learned from the seen classes, and robustifies the adapted VLM to unseen categories. Our empirical evaluation on various datasets validates the effectiveness of Fed-MP.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01232
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open-Vocabulary Federated Learning with Multimodal Prototyping
Zeng, Huimin
Yue, Zhenrui
Wang, Dong
Computation and Language
Computer Vision and Pattern Recognition
Existing federated learning (FL) studies usually assume the training label space and test label space are identical. However, in real-world applications, this assumption is too ideal to be true. A new user could come up with queries that involve data from unseen classes, and such open-vocabulary queries would directly defect such FL systems. Therefore, in this work, we explicitly focus on the under-explored open-vocabulary challenge in FL. That is, for a new user, the global server shall understand her/his query that involves arbitrary unknown classes. To address this problem, we leverage the pre-trained vision-language models (VLMs). In particular, we present a novel adaptation framework tailored for VLMs in the context of FL, named as Federated Multimodal Prototyping (Fed-MP). Fed-MP adaptively aggregates the local model weights based on light-weight client residuals, and makes predictions based on a novel multimodal prototyping mechanism. Fed-MP exploits the knowledge learned from the seen classes, and robustifies the adapted VLM to unseen categories. Our empirical evaluation on various datasets validates the effectiveness of Fed-MP.
title Open-Vocabulary Federated Learning with Multimodal Prototyping
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.01232