Accelerated Training of Federated Learning via Second-Order Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sen, Mrinmay, Nair, Sidhant R, Mohan, C Krishna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912402527748096
author Sen, Mrinmay
Nair, Sidhant R
Mohan, C Krishna
author_facet Sen, Mrinmay
Nair, Sidhant R
Mohan, C Krishna
contents This paper explores second-order optimization methods in Federated Learning (FL), addressing the critical challenges of slow convergence and the excessive communication rounds required to achieve optimal performance from the global model. While existing surveys in FL primarily focus on challenges related to statistical and device label heterogeneity, as well as privacy and security concerns in first-order FL methods, less attention has been given to the issue of slow model training. This slow training often leads to the need for excessive communication rounds or increased communication costs, particularly when data across clients are highly heterogeneous. In this paper, we examine various FL methods that leverage second-order optimization to accelerate the training process. We provide a comprehensive categorization of state-of-the-art second-order FL methods and compare their performance based on convergence speed, computational cost, memory usage, transmission overhead, and generalization of the global model. Our findings show the potential of incorporating Hessian curvature through second-order optimization into FL and highlight key challenges, such as the efficient utilization of Hessian and its inverse in FL. This work lays the groundwork for future research aimed at developing scalable and efficient federated optimization methods for improving the training of the global model in FL.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23588
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerated Training of Federated Learning via Second-Order Methods
Sen, Mrinmay
Nair, Sidhant R
Mohan, C Krishna
Machine Learning
Distributed, Parallel, and Cluster Computing
68Q25, 68T05, 90C06, 90C25, 90C30
I.2.6; G.1.6; C.2.4; C.4
This paper explores second-order optimization methods in Federated Learning (FL), addressing the critical challenges of slow convergence and the excessive communication rounds required to achieve optimal performance from the global model. While existing surveys in FL primarily focus on challenges related to statistical and device label heterogeneity, as well as privacy and security concerns in first-order FL methods, less attention has been given to the issue of slow model training. This slow training often leads to the need for excessive communication rounds or increased communication costs, particularly when data across clients are highly heterogeneous. In this paper, we examine various FL methods that leverage second-order optimization to accelerate the training process. We provide a comprehensive categorization of state-of-the-art second-order FL methods and compare their performance based on convergence speed, computational cost, memory usage, transmission overhead, and generalization of the global model. Our findings show the potential of incorporating Hessian curvature through second-order optimization into FL and highlight key challenges, such as the efficient utilization of Hessian and its inverse in FL. This work lays the groundwork for future research aimed at developing scalable and efficient federated optimization methods for improving the training of the global model in FL.
title Accelerated Training of Federated Learning via Second-Order Methods
topic Machine Learning
Distributed, Parallel, and Cluster Computing
68Q25, 68T05, 90C06, 90C25, 90C30
I.2.6; G.1.6; C.2.4; C.4
url https://arxiv.org/abs/2505.23588