Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917360686858240 |
|---|---|
| author | Guo, Xin Zhao, Chunrui Jia, Hong Dang, Ting Huang, Gongping Zheng, Xianrui Gao, Yan |
| author_facet | Guo, Xin Zhao, Chunrui Jia, Hong Dang, Ting Huang, Gongping Zheng, Xianrui Gao, Yan |
| contents | Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant heterogeneity: clients differ in computational capacity, causing straggler effects under unified fine-tuning, while diverse downstream tasks require different representation depths, making full-model updates inefficient. To address these challenges, we propose an adaptive federated fine-tuning framework with early exits. Lightweight prediction heads are inserted at intermediate layers of the SSL backbone, allowing clients to terminate computation based on local constraints and task requirements. We further introduce a layer-wise, depth-aware partial aggregation strategy to better utilize representations from different network depths. Experiments show that the framework reduces edge overhead, supports heterogeneous hardware, and maintains competitive performance in resource-constrained federated environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_21888 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations Guo, Xin Zhao, Chunrui Jia, Hong Dang, Ting Huang, Gongping Zheng, Xianrui Gao, Yan Audio and Speech Processing Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant heterogeneity: clients differ in computational capacity, causing straggler effects under unified fine-tuning, while diverse downstream tasks require different representation depths, making full-model updates inefficient. To address these challenges, we propose an adaptive federated fine-tuning framework with early exits. Lightweight prediction heads are inserted at intermediate layers of the SSL backbone, allowing clients to terminate computation based on local constraints and task requirements. We further introduce a layer-wise, depth-aware partial aggregation strategy to better utilize representations from different network depths. Experiments show that the framework reduces edge overhead, supports heterogeneous hardware, and maintains competitive performance in resource-constrained federated environments. |
| title | Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2603.21888 |