Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Xin, Zhao, Chunrui, Jia, Hong, Dang, Ting, Huang, Gongping, Zheng, Xianrui, Gao, Yan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917360686858240
author Guo, Xin
Zhao, Chunrui
Jia, Hong
Dang, Ting
Huang, Gongping
Zheng, Xianrui
Gao, Yan
author_facet Guo, Xin
Zhao, Chunrui
Jia, Hong
Dang, Ting
Huang, Gongping
Zheng, Xianrui
Gao, Yan
contents Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant heterogeneity: clients differ in computational capacity, causing straggler effects under unified fine-tuning, while diverse downstream tasks require different representation depths, making full-model updates inefficient. To address these challenges, we propose an adaptive federated fine-tuning framework with early exits. Lightweight prediction heads are inserted at intermediate layers of the SSL backbone, allowing clients to terminate computation based on local constraints and task requirements. We further introduce a layer-wise, depth-aware partial aggregation strategy to better utilize representations from different network depths. Experiments show that the framework reduces edge overhead, supports heterogeneous hardware, and maintains competitive performance in resource-constrained federated environments.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
Guo, Xin
Zhao, Chunrui
Jia, Hong
Dang, Ting
Huang, Gongping
Zheng, Xianrui
Gao, Yan
Audio and Speech Processing
Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant heterogeneity: clients differ in computational capacity, causing straggler effects under unified fine-tuning, while diverse downstream tasks require different representation depths, making full-model updates inefficient. To address these challenges, we propose an adaptive federated fine-tuning framework with early exits. Lightweight prediction heads are inserted at intermediate layers of the SSL backbone, allowing clients to terminate computation based on local constraints and task requirements. We further introduce a layer-wise, depth-aware partial aggregation strategy to better utilize representations from different network depths. Experiments show that the framework reduces edge overhead, supports heterogeneous hardware, and maintains competitive performance in resource-constrained federated environments.
title Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
topic Audio and Speech Processing
url https://arxiv.org/abs/2603.21888