Federated Learning Framework for Scalable AI in Heterogeneous HPC and Cloud Environments

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ghimire, Sangam, Timalsina, Paribartan, Bhurtel, Nirjal, Neupane, Bishal, Shrestha, Bigyan Byanju, Bhattarai, Subarna, Gaire, Prajwal, Thapa, Jessica, Jha, Sudan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917103404056576
author Ghimire, Sangam
Timalsina, Paribartan
Bhurtel, Nirjal
Neupane, Bishal
Shrestha, Bigyan Byanju
Bhattarai, Subarna
Gaire, Prajwal
Thapa, Jessica
Jha, Sudan
author_facet Ghimire, Sangam
Timalsina, Paribartan
Bhurtel, Nirjal
Neupane, Bishal
Shrestha, Bigyan Byanju
Bhattarai, Subarna
Gaire, Prajwal
Thapa, Jessica
Jha, Sudan
contents As the demand grows for scalable and privacy-aware AI systems, Federated Learning (FL) has emerged as a promising solution, allowing decentralized model training without moving raw data. At the same time, the combination of high-performance computing (HPC) and cloud infrastructure offers vast computing power but introduces new complexities, especially when dealing with heterogeneous hardware, communication limits, and non-uniform data. In this work, we present a federated learning framework built to run efficiently across mixed HPC and cloud environments. Our system addresses key challenges such as system heterogeneity, communication overhead, and resource scheduling, while maintaining model accuracy and data privacy. Through experiments on a hybrid testbed, we demonstrate strong performance in terms of scalability, fault tolerance, and convergence, even under non-Independent and Identically Distributed (non-IID) data distributions and varied hardware. These results highlight the potential of federated learning as a practical approach to building scalable Artificial Intelligence (AI) systems in modern, distributed computing settings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Federated Learning Framework for Scalable AI in Heterogeneous HPC and Cloud Environments
Ghimire, Sangam
Timalsina, Paribartan
Bhurtel, Nirjal
Neupane, Bishal
Shrestha, Bigyan Byanju
Bhattarai, Subarna
Gaire, Prajwal
Thapa, Jessica
Jha, Sudan
Distributed, Parallel, and Cluster Computing
Machine Learning
As the demand grows for scalable and privacy-aware AI systems, Federated Learning (FL) has emerged as a promising solution, allowing decentralized model training without moving raw data. At the same time, the combination of high-performance computing (HPC) and cloud infrastructure offers vast computing power but introduces new complexities, especially when dealing with heterogeneous hardware, communication limits, and non-uniform data. In this work, we present a federated learning framework built to run efficiently across mixed HPC and cloud environments. Our system addresses key challenges such as system heterogeneity, communication overhead, and resource scheduling, while maintaining model accuracy and data privacy. Through experiments on a hybrid testbed, we demonstrate strong performance in terms of scalability, fault tolerance, and convergence, even under non-Independent and Identically Distributed (non-IID) data distributions and varied hardware. These results highlight the potential of federated learning as a practical approach to building scalable Artificial Intelligence (AI) systems in modern, distributed computing settings.
title Federated Learning Framework for Scalable AI in Heterogeneous HPC and Cloud Environments
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2511.19479