Saved in:
Bibliographic Details
Main Authors: Zhang, Yu, Du, Qingfeng, Lv, Jiaqi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.12025
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916693216854016
author Zhang, Yu
Du, Qingfeng
Lv, Jiaqi
author_facet Zhang, Yu
Du, Qingfeng
Lv, Jiaqi
contents Federated Learning (FL) enables decentralized model training across multiple parties while preserving privacy. However, most FL systems assume clients hold only unimodal data, limiting their real-world applicability, as institutions often possess multimodal data. Moreover, the lack of labeled data further constrains the performance of most FL methods. In this work, we propose FedEPA, a novel FL framework for multimodal learning. FedEPA employs a personalized local model aggregation strategy that leverages labeled data on clients to learn personalized aggregation weights, thereby alleviating the impact of data heterogeneity. We also propose an unsupervised modality alignment strategy that works effectively with limited labeled data. Specifically, we decompose multimodal features into aligned features and context features. We then employ contrastive learning to align the aligned features across modalities, ensure the independence between aligned features and context features within each modality, and promote the diversity of context features. A multimodal feature fusion strategy is introduced to obtain a joint embedding. The experimental results show that FedEPA significantly outperforms existing FL methods in multimodal classification tasks under limited labeled data conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning
Zhang, Yu
Du, Qingfeng
Lv, Jiaqi
Machine Learning
Federated Learning (FL) enables decentralized model training across multiple parties while preserving privacy. However, most FL systems assume clients hold only unimodal data, limiting their real-world applicability, as institutions often possess multimodal data. Moreover, the lack of labeled data further constrains the performance of most FL methods. In this work, we propose FedEPA, a novel FL framework for multimodal learning. FedEPA employs a personalized local model aggregation strategy that leverages labeled data on clients to learn personalized aggregation weights, thereby alleviating the impact of data heterogeneity. We also propose an unsupervised modality alignment strategy that works effectively with limited labeled data. Specifically, we decompose multimodal features into aligned features and context features. We then employ contrastive learning to align the aligned features across modalities, ensure the independence between aligned features and context features within each modality, and promote the diversity of context features. A multimodal feature fusion strategy is introduced to obtain a joint embedding. The experimental results show that FedEPA significantly outperforms existing FL methods in multimodal classification tasks under limited labeled data conditions.
title FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning
topic Machine Learning
url https://arxiv.org/abs/2504.12025