TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Guanxiong, Mirmehdi, Majid, Abdallah, Zahraa, Santos-Rodriguez, Raul, Craddock, Ian, Filho, Telmo de Menezes e Silva
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916960380387328
author Sun, Guanxiong
Mirmehdi, Majid
Abdallah, Zahraa
Santos-Rodriguez, Raul
Craddock, Ian
Filho, Telmo de Menezes e Silva
author_facet Sun, Guanxiong
Mirmehdi, Majid
Abdallah, Zahraa
Santos-Rodriguez, Raul
Craddock, Ian
Filho, Telmo de Menezes e Silva
contents Real-world federated learning faces two key challenges: limited access to labelled data and the presence of heterogeneous multi-modal inputs. This paper proposes TACTFL, a unified framework for semi-supervised multi-modal federated learning. TACTFL introduces a modality-agnostic temporal contrastive training scheme that conducts representation learning from unlabelled client data by leveraging temporal alignment across modalities. However, as clients perform self-supervised training on heterogeneous data, local models may diverge semantically. To mitigate this, TACTFL incorporates a similarity-guided model aggregation strategy that dynamically weights client models based on their representational consistency, promoting global alignment. Extensive experiments across diverse benchmarks and modalities, including video, audio, and wearable sensors, demonstrate that TACTFL achieves state-of-the-art performance. For instance, on the UCF101 dataset with only 10% labelled data, TACTFL attains 68.48% top-1 accuracy, significantly outperforming the FedOpt baseline of 35.35%. Code will be released upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
Sun, Guanxiong
Mirmehdi, Majid
Abdallah, Zahraa
Santos-Rodriguez, Raul
Craddock, Ian
Filho, Telmo de Menezes e Silva
Distributed, Parallel, and Cluster Computing
Real-world federated learning faces two key challenges: limited access to labelled data and the presence of heterogeneous multi-modal inputs. This paper proposes TACTFL, a unified framework for semi-supervised multi-modal federated learning. TACTFL introduces a modality-agnostic temporal contrastive training scheme that conducts representation learning from unlabelled client data by leveraging temporal alignment across modalities. However, as clients perform self-supervised training on heterogeneous data, local models may diverge semantically. To mitigate this, TACTFL incorporates a similarity-guided model aggregation strategy that dynamically weights client models based on their representational consistency, promoting global alignment. Extensive experiments across diverse benchmarks and modalities, including video, audio, and wearable sensors, demonstrate that TACTFL achieves state-of-the-art performance. For instance, on the UCF101 dataset with only 10% labelled data, TACTFL attains 68.48% top-1 accuracy, significantly outperforming the FedOpt baseline of 35.35%. Code will be released upon publication.
title TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.17532