Adaptive Federated Learning Over the Air

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Chenhao, Chen, Zihan, Pappas, Nikolaos, Yang, Howard H., Quek, Tony Q. S., Poor, H. Vincent
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917610370629632
author Wang, Chenhao
Chen, Zihan
Pappas, Nikolaos
Yang, Howard H.
Quek, Tony Q. S.
Poor, H. Vincent
author_facet Wang, Chenhao
Chen, Zihan
Pappas, Nikolaos
Yang, Howard H.
Quek, Tony Q. S.
Poor, H. Vincent
contents We propose a federated version of adaptive gradient methods, particularly AdaGrad and Adam, within the framework of over-the-air model training. This approach capitalizes on the inherent superposition property of wireless channels, facilitating fast and scalable parameter aggregation. Meanwhile, it enhances the robustness of the model training process by dynamically adjusting the stepsize in accordance with the global gradient update. We derive the convergence rate of the training algorithms, encompassing the effects of channel fading and interference, for a broad spectrum of nonconvex loss functions. Our analysis shows that the AdaGrad-based algorithm converges to a stationary point at the rate of $\mathcal{O}( \ln{(T)} /{ T^{ 1 - \frac{1}α } } )$, where $α$ represents the tail index of the electromagnetic interference. This result indicates that the level of heavy-tailedness in interference distribution plays a crucial role in the training efficiency: the heavier the tail, the slower the algorithm converges. In contrast, an Adam-like algorithm converges at the $\mathcal{O}( 1/T )$ rate, demonstrating its advantage in expediting the model training process. We conduct extensive experiments that corroborate our theoretical findings and affirm the practical efficacy of our proposed federated adaptive gradient methods.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06528
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Federated Learning Over the Air
Wang, Chenhao
Chen, Zihan
Pappas, Nikolaos
Yang, Howard H.
Quek, Tony Q. S.
Poor, H. Vincent
Machine Learning
Information Theory
Networking and Internet Architecture
We propose a federated version of adaptive gradient methods, particularly AdaGrad and Adam, within the framework of over-the-air model training. This approach capitalizes on the inherent superposition property of wireless channels, facilitating fast and scalable parameter aggregation. Meanwhile, it enhances the robustness of the model training process by dynamically adjusting the stepsize in accordance with the global gradient update. We derive the convergence rate of the training algorithms, encompassing the effects of channel fading and interference, for a broad spectrum of nonconvex loss functions. Our analysis shows that the AdaGrad-based algorithm converges to a stationary point at the rate of $\mathcal{O}( \ln{(T)} /{ T^{ 1 - \frac{1}α } } )$, where $α$ represents the tail index of the electromagnetic interference. This result indicates that the level of heavy-tailedness in interference distribution plays a crucial role in the training efficiency: the heavier the tail, the slower the algorithm converges. In contrast, an Adam-like algorithm converges at the $\mathcal{O}( 1/T )$ rate, demonstrating its advantage in expediting the model training process. We conduct extensive experiments that corroborate our theoretical findings and affirm the practical efficacy of our proposed federated adaptive gradient methods.
title Adaptive Federated Learning Over the Air
topic Machine Learning
Information Theory
Networking and Internet Architecture
url https://arxiv.org/abs/2403.06528