Modality Alignment Meets Federated Broadcasting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Yuting, Tang, Shengeng, Xu, Xiaohua, Cheng, Lechao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910712577654784
author Ma, Yuting
Tang, Shengeng
Xu, Xiaohua
Cheng, Lechao
author_facet Ma, Yuting
Tang, Shengeng
Xu, Xiaohua
Cheng, Lechao
contents Federated learning (FL) has emerged as a powerful approach to safeguard data privacy by training models across distributed edge devices without centralizing local data. Despite advancements in homogeneous data scenarios, maintaining performance between the global and local clients in FL over heterogeneous data remains challenging due to data distribution variations that degrade model convergence and increase computational costs. This paper introduces a novel FL framework leveraging modality alignment, where a text encoder resides on the server, and image encoders operate on local devices. Inspired by multi-modal learning paradigms like CLIP, this design aligns cross-client learning by treating server-client communications akin to multi-modal broadcasting. We initialize with a pre-trained model to mitigate overfitting, updating select parameters through low-rank adaptation (LoRA) to meet computational demand and performance efficiency. Local models train independently and communicate updates to the server, which aggregates parameters via a query-based method, facilitating cross-client knowledge sharing and performance improvement under extreme heterogeneity. Extensive experiments on benchmark datasets demonstrate the efficacy in maintaining generalization and robustness, even in highly heterogeneous settings.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15837
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Modality Alignment Meets Federated Broadcasting
Ma, Yuting
Tang, Shengeng
Xu, Xiaohua
Cheng, Lechao
Computer Vision and Pattern Recognition
Federated learning (FL) has emerged as a powerful approach to safeguard data privacy by training models across distributed edge devices without centralizing local data. Despite advancements in homogeneous data scenarios, maintaining performance between the global and local clients in FL over heterogeneous data remains challenging due to data distribution variations that degrade model convergence and increase computational costs. This paper introduces a novel FL framework leveraging modality alignment, where a text encoder resides on the server, and image encoders operate on local devices. Inspired by multi-modal learning paradigms like CLIP, this design aligns cross-client learning by treating server-client communications akin to multi-modal broadcasting. We initialize with a pre-trained model to mitigate overfitting, updating select parameters through low-rank adaptation (LoRA) to meet computational demand and performance efficiency. Local models train independently and communicate updates to the server, which aggregates parameters via a query-based method, facilitating cross-client knowledge sharing and performance improvement under extreme heterogeneity. Extensive experiments on benchmark datasets demonstrate the efficacy in maintaining generalization and robustness, even in highly heterogeneous settings.
title Modality Alignment Meets Federated Broadcasting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15837